Why This Role Stands Out
This hybrid role offers a fantastic opportunity to lead critical incident response and drive impactful automation initiatives within a reputable technology company. You'll thrive here if you are a seasoned DevOps leader with a passion for problem-solving, continuous improvement, and mentoring a talented team. Embrace this chance to shape operational excellence and advance your career in a supportive environment.
Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Chicago, IL, United States
Posted
3 weeks ago
SplunkApacheDatadogJenkinsKafkaKubernetesVault
Job Description
***We are unable to sponsor as this is a permanent full-time role***
***Hybrid, 3 days onsite, 2 days remote***
Responsibilities:
- Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
- Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams.
- Serve as Tier 3 escalation point for complex incidents beyond L2 capability triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.
- Drive post-incident reviews for all P1 and P2 incidents, ensuring root cause is identified, documented, and actioned not filed.
- Define, publish, and enforce SLA targets across all severity levels
- Monitor SLA compliance in real time; escalate breaches immediately and report trends to leadership on a sprint cadence.
- Publish monthly SLA compliance reports to leadership with trend analysis and improvement actions.
- Lead alert tuning and noise reduction initiatives across monitoring toolsets on-call engineers are paged for situations requiring human judgement, not system noise.
- Track and publish the alert-to-incident ratio each sprint; hold the team accountable to a visible and improving trend.
- Lead automation and tooling initiatives to reduce toil, accelerate triage, and eliminate manual steps from the support workflow.
- Ensure accurate, complete documentation for every incident symptoms, steps taken, diagnostics, resolution, and RCA where applicable.
- Own the runbook library every novel resolution produces a runbook published to L1 before the incident is closed; coverage gaps are tracked and closed sprint-on-sprint.
- Lead a team of 6 10 L1 and L2 support engineers and contingent labor within the Environment Operations function.
- Manage team scheduling to ensure full coverage of production support windows including on-call rotations, shift handoffs, and escalation availability for 247 support responsibilities.
Qualifications:
- Minimum 5 years of hands-on environment operations, production support, or infrastructure operations experience, including interdisciplinary experience across four or more of the following: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation.
- Technical experience and comprehensive knowledge of production environment failure modes including deployment failures, configuration drift, platform instability, and integration breakdowns and the methodologies used to diagnose and resolve them.
- Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent).
- Shift work and on-call availability required including 247 on-call response capacity and availability during planned and emergency maintenance windows.
- Previous people management or team lead experience required; formal people management experience strongly preferred.
- Deployment & Pipeline tooling: Harness (continuous delivery pipelines, deployment verification, rollback automation), Jenkins (CI/CD pipeline management, job configuration, build troubleshooting), GitHub (branching strategies, pull request workflows, pipeline integration).
- Container & orchestration platforms: Kubernetes (k8s) pod lifecycle management, namespace operations, log retrieval, resource troubleshooting, and coordination with Platform teams on cluster-level issues.
- Messaging & streaming platforms: Apache Kafka topic management, consumer group monitoring, lag analysis, and escalation to Platform for broker-level issues.
- Secrets & configuration management: HashiCorp Vault secrets retrieval, token/lease troubleshooting, policy review, and escalation to Security teams for certificate and secrets rotation.
- Monitoring & observability: Proficiency in at least two production monitoring toolsets (e.g. Splunk, Dynatrace, Datadog, AppDynamics, PrometheGrafana) alert triage, dashboard interpretation, log analysis, and tuning requests.
- Middleware platforms: Working knowledge of middleware infrastructure including application servers, messaging brokers, storage integrations, and network-layer dependencies sufficient to triage, gather diagnostics, and route correctly to L3.
Similar jobs
- TE
Senior Cloud DevOps Engineer – Platform Engineering - Atlanta, GA, Charlotte, NC, Tampa, FL, Raleigh, NC, Durham, NC, Orlando, FL, Columbia, SC.
NewTechniPros, LLC
Atlanta, GA🇺🇸Hybrid16 hours agoDockerAWSArgoCD+11Technology - TE
DevOps Engineer - New York City, NY, Boston, MA, Hartford, CT, Princeton, NJ, Newark, NJ.
NewTechniPros, LLC
New York, NY🇺🇸Hybrid16 hours agoDockerAWSMLOps+11Technology - HP
Devops Azure Engineer (W2 Contract)
NewHPTech Inc.
Plano, TX🇺🇸Hybrid16 hours agoDockerAWSAzure+3Technology - AC
Lead Devops SRE Engineer
NewAlltech Consulting Services, Inc.
Jersey City, NJ🇺🇸On-site16 hours agoDockerAWSELK+12Technology - HP
Azure DevOps / Automation Engineer
NewHPTech Inc.
Plano, TX🇺🇸Hybrid16 hours agoDockerAWSAzure+3Technology - TR
ELK / Stash SRE Engineer (DevOps), Eden Prairie, MN (Onsite & Locals Only)
NewTror
Eden Prairie, MN🇺🇸On-site16 hours agoAWSELKLogstash+3Technology