Haystack
← Back to Jobs
Technology
AC

Lead Devops SRE Engineer

Alltech Consulting Services, Inc.Jersey City, NJ🇺🇸United StatesPosted Sep 17, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Jersey City, NJ, United States
Posted
17 hours ago
DockerAWSELKAzureBashDatadogGitHub ActionsGitLab CIGrafanaJavaJenkinsKubernetesPrometheusPythonTerraform

Job Description

Lead Devops SRE Engineer
Location: Jersey City, NJ/Onsite
Job Summary
We are seeking an experienced Lead Site Reliability Engineer (SRE) to lead the reliability, scalability, automation, and operational excellence of enterprise production platforms. The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, CI/CD, observability, and production incident management, along with the ability to mentor and guide SRE/DevOps engineers.
Key Responsibilities
  • Lead SRE initiatives focused on system reliability, availability, scalability, and performance.
  • Design and maintain highly available and resilient AWS cloud infrastructure.
  • Manage and troubleshoot production Kubernetes/EKS environments.
  • Define and implement SLIs, SLOs, SLAs, and error budgets.
  • Lead P1/P2 incident response, troubleshooting, RCA, and post-incident reviews.
  • Develop automation to eliminate manual operational tasks and reduce operational toil.
  • Build and maintain Infrastructure as Code (IaC) using Terraform.
  • Design and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
  • Implement and improve monitoring, logging, alerting, and observability.
  • Work with tools such as Prometheus, Grafana, CloudWatch, ELK, Datadog, and OpenTelemetry.
  • Develop automation and operational tools using Python, Bash, or Go.
  • Drive capacity planning, performance optimization, disaster recovery, and business continuity initiatives.
  • Establish reliability engineering best practices across development and infrastructure teams.
  • Collaborate with application, cloud, security, and architecture teams to improve production readiness.
  • Lead technical discussions, architecture reviews, and reliability improvements.
  • Mentor junior and senior SRE/DevOps engineers and provide technical guidance.
  • Participate in on-call rotations and ensure production systems meet defined reliability objectives.
Required Skills
  • 7+ years of experience in SRE, DevOps, Cloud Engineering, or Platform Engineering.
  • Strong hands-on experience with AWS.
  • Strong expertise in Kubernetes/EKS and Docker.
  • Strong experience with Terraform / Infrastructure as Code.
  • Hands-on experience with CI/CD pipelines.
  • Strong knowledge of Linux administration and troubleshooting.
  • Experience with Prometheus, Grafana, CloudWatch, ELK, Datadog, or similar observability tools.
  • Strong scripting/programming experience with Python, Bash, or Go.
  • Experience with production incident management, RCA, postmortems, and on-call operations.
  • Understanding of SLI, SLO, SLA, error budgets, MTTR, and MTBF.
  • Experience with cloud networking, security, IAM, and high-availability architectures.
  • Strong communication, problem-solving, and leadership skills.

Similar jobs