Why This Role Stands Out
This Manager, Software Engineering DevOps role offers a compelling opportunity to lead and mentor a talented team, driving critical incident management and release engineering within a reputable company. If you excel in Kubernetes, CI/CD pipelines, and possess strong leadership skills, you'll thrive in this hybrid position with a competitive salary and bonus structure. Apply now to shape the future of DevOps at Request Technology, LLC!
Quick Overview
Job Description
NO SPONSORSHIP - NO OPT
Manager, Software Engineering DevOps
Salary:$170k - $200k plus 15% Bonus
LOCATION: Chicago, IL
Hybrid 3 days onsite and 2 days remote
You will lead a team of 6-10, 24/7 support, L1, L2 support. Application deployment container platform ops middleware support incident management monitoring observability configuration release engineering platform operations scripting automation SLA frameworks harness Jenkins Kubernetes apache kakfa splunk Dynatrace datadog.
- Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
- Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams.
- Serve as Tier 3 escalation point for complex incidents beyond L2 capability triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.
- Own the full incident lifecycle from first alert through to RCA documentation and permanent fix or accepted workaround.
- Drive post-incident reviews for all P1 and P2 incidents, ensuring root cause is identified, documented, and actioned not filed.
Technical Skills:
- Deployment & Pipeline tooling: Harness (continuous delivery pipelines, deployment verification, rollback automation), Jenkins (CI/CD pipeline management, job configuration, build troubleshooting), GitHub (branching strategies, pull request workflows, pipeline integration).
- Container & orchestration platforms: Kubernetes (k8s) pod lifecycle management, namespace operations, log retrieval, resource troubleshooting, and coordination with Platform teams on cluster-level issues.
- Messaging & streaming platforms: Apache Kafka topic management, consumer group monitoring, lag analysis, and escalation to Platform for broker-level issues.
- Secrets & configuration management: HashiCorp Vault secrets retrieval, token/lease troubleshooting, policy review, and escalation to Security teams for certificate and secrets rotation.
- Monitoring & observability: Proficiency in at least two production monitoring toolsets (e.g. Splunk, Dynatrace, Datadog, AppDynamics, PrometheGrafana) alert triage, dashboard interpretation, log analysis, and tuning requests.
- Middleware platforms: Working knowledge of middleware infrastructure including application servers, messaging brokers, storage integrations, and network-layer dependencies sufficient to triage, gather diagnostics, and route correctly to L3.
- Incident and ticketing platforms: ServiceNow or equivalent ITSM tooling incident creation, SLA tracking, problem record management, and reporting.
- MTTR and operational metrics: Ability to build and maintain operational dashboards and reports covering MTTR, SLA compliance, alert-to-incident ratio, repeat incident rate, and deployment success rate.
Education and/or Experience:
- Minimum 5 years of hands-on environment operations, production support, or infrastructure operations experience, including interdisciplinary experience across four or more of the following: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation.
- Technical experience and comprehensive knowledge of production environment failure modes including deployment failures, configuration drift, platform instability, and integration breakdowns and the methodologies used to diagnose and resolve them.
- Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent).
- Familiarity with financial services or other regulated-industry production environments is a strong advantage understanding of change governance, audit requirements, and production access controls.
- Industry knowledge of current and emerging practices in environment operations, platform reliability, and support automation.
- Shift work and on-call availability required including 247 on-call response capacity and availability during planned and emergency maintenance windows.
- Previous people management or team lead experience required; formal people management experience
Similar jobs
- GL
DevOps Engineer - Plano, TX - W2 Position ONLY
NewGeopaq Logic
Plano, TX🇺🇸HybridYesterdayLLMTechnology - IN
Systems Operations Engineer
NewInnova
Charlotte, NC🇺🇸$40 - $45/hrHybridYesterdayAWSTechnology - CY
Site Reliability Engineer
NewcyberThink, Inc.
Austin, TX🇺🇸HybridYesterdayAWSSplunkAnsible+8Technology - JM
Infrastructure Engineer III - Site Reliability
NewJ.P. Morgan
Plano, Texas🇺🇸On-site1 hour agoGCPAWSSplunk+6Technology - OR
Senior Manager, Site Reliability Engineering
NewOracle
Reston, Virginia🇺🇸$121.5k - $264.1k/yrHybrid1 hour agoOracleTechnology - BA
SRE Virtual Desktop Operations Engineer - AVP
NewBarclays
New York City, New York🇺🇸Hybrid1 hour agoSplunkActive DirectoryAnsible+7Technology