Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
Jersey City, NJ, United States
Posted
17 hours ago
DockerAWSELKAzureBashDatadogGitHub ActionsGitLab CIGrafanaJavaJenkinsKubernetesPrometheusPythonTerraform
Job Description
Lead Devops SRE Engineer
Location: Jersey City, NJ/Onsite
Job Summary
We are seeking an experienced Lead Site Reliability Engineer (SRE) to lead the reliability, scalability, automation, and operational excellence of enterprise production platforms. The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, CI/CD, observability, and production incident management, along with the ability to mentor and guide SRE/DevOps engineers.
Key Responsibilities
- Lead SRE initiatives focused on system reliability, availability, scalability, and performance.
- Design and maintain highly available and resilient AWS cloud infrastructure.
- Manage and troubleshoot production Kubernetes/EKS environments.
- Define and implement SLIs, SLOs, SLAs, and error budgets.
- Lead P1/P2 incident response, troubleshooting, RCA, and post-incident reviews.
- Develop automation to eliminate manual operational tasks and reduce operational toil.
- Build and maintain Infrastructure as Code (IaC) using Terraform.
- Design and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Implement and improve monitoring, logging, alerting, and observability.
- Work with tools such as Prometheus, Grafana, CloudWatch, ELK, Datadog, and OpenTelemetry.
- Develop automation and operational tools using Python, Bash, or Go.
- Drive capacity planning, performance optimization, disaster recovery, and business continuity initiatives.
- Establish reliability engineering best practices across development and infrastructure teams.
- Collaborate with application, cloud, security, and architecture teams to improve production readiness.
- Lead technical discussions, architecture reviews, and reliability improvements.
- Mentor junior and senior SRE/DevOps engineers and provide technical guidance.
- Participate in on-call rotations and ensure production systems meet defined reliability objectives.
Required Skills
- 7+ years of experience in SRE, DevOps, Cloud Engineering, or Platform Engineering.
- Strong hands-on experience with AWS.
- Good Experience in Java
- Strong expertise in Kubernetes/EKS and Docker.
- Strong experience with Terraform / Infrastructure as Code.
- Hands-on experience with CI/CD pipelines.
- Strong knowledge of Linux administration and troubleshooting.
- Experience with Prometheus, Grafana, CloudWatch, ELK, Datadog, or similar observability tools.
- Strong scripting/programming experience with Python, Bash, or Go.
- Experience with production incident management, RCA, postmortems, and on-call operations.
- Understanding of SLI, SLO, SLA, error budgets, MTTR, and MTBF.
- Experience with cloud networking, security, IAM, and high-availability architectures.
- Strong communication, problem-solving, and leadership skills.
Similar jobs
- IN
Site Reliability Engineer (SRE) Lead - Charlotte, NC
NewInfojini
Charlotte, NC🇺🇸On-site17 hours agoSplunkTechnology - PG
Sr. Site Reliability Engineer - SRE, Onsite - 70170
NewPRIMUS Global Services Inc.
TX🇺🇸On-site17 hours agoMongoDBOracleSQL+11Technology - PR
Full-Time - 10+ Principal Site Reliability Engineer - Atlanta, GA (Hybrid)
NewProhires
Atlanta, GA🇺🇸Hybrid17 hours agoOracleSQLSQL Server+12Technology - IS
DevOps Engineer
NewInnova Solutions, Inc
Dallas, TX🇺🇸$50 - $55/hrOn-site17 hours agoAWSSonarQubeDatadog+3Technology - LM
Technical Business Analyst with Azure and Devops : Houston, TX (one day a week in the office)
NewLightning Minds Inc.
Houston, TX🇺🇸Hybrid17 hours agoAzureTechnology - AT
AWS Cloud Platform Engineer / Senior AWS Infrastructure Engineer at New Jersey
NewArkhya Tech
Woodbridge Township, NJ🇺🇸On-site17 hours agoAWSSSOAnsible+6Technology