← Back to Jobs
Engineering
AT
Major Incident Engineer
Apidel TechnologiesCincinnati, OH🇺🇸United StatesPosted 24 Jul 2026
Quick Overview
Work Type
On Site
Level
Mid Senior
Job Description
*]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))] dir=auto data-turn-id=request-6a627f82-95dc-83ee-8a48-28e9a775e98b-9 data-turn-id-container=request-6a627f82-95dc-83ee-8a48-28e9a775e98b-9 data-testid=conversation-turn-36 data-turn=assistant>
Location: Cincinnati, OH (100% Onsite)
Travel: Occasional travel required
On-Call: Participation in a rotating on-call schedule required
Job Overview:
o We are seeking a Major Incident & Site Reliability Engineer to support critical production environments by leading high-severity incident response, driving service restoration, and improving system reliability.
o This role requires strong technical troubleshooting skills, cross-functional collaboration, and experience managing enterprise production incidents in a 24x7 environment.
Responsibilities:
o Lead P1/P2 major incident response, coordinating cross-functional teams to restore critical services.
o Facilitate incident bridge calls, provide executive communications, and drive timely resolution.
o Perform Root Cause Analysis (RCA) and partner with engineering teams to implement corrective actions.
o Troubleshoot production issues across applications, infrastructure, databases, and cloud environments.
o Support production operations, monitoring, deployments, and operational readiness.
o Participate in on-call rotations and collaborate with development, infrastructure, and operations teams to improve platform reliability.
Required Skills:
o Experience leading Major Incident Management (P1/P2) in a 24x7 enterprise environment.
o Strong background in Site Reliability Engineering (SRE) or L3 Production Support.
o Experience with ServiceNow, ITIL, and enterprise monitoring tools such as Splunk, Grafana, Dynatrace, or similar.
o Hands-on troubleshooting experience with Linux/Unix, cloud platforms, databases, and distributed applications.
o Scripting experience with Python, Bash, or PowerShell is preferred.
o Excellent communication skills with the ability to coordinate technical and business stakeholders during critical incidents.
Travel: Occasional travel required
On-Call: Participation in a rotating on-call schedule required
Job Overview:
o We are seeking a Major Incident & Site Reliability Engineer to support critical production environments by leading high-severity incident response, driving service restoration, and improving system reliability.
o This role requires strong technical troubleshooting skills, cross-functional collaboration, and experience managing enterprise production incidents in a 24x7 environment.
Responsibilities:
o Lead P1/P2 major incident response, coordinating cross-functional teams to restore critical services.
o Facilitate incident bridge calls, provide executive communications, and drive timely resolution.
o Perform Root Cause Analysis (RCA) and partner with engineering teams to implement corrective actions.
o Troubleshoot production issues across applications, infrastructure, databases, and cloud environments.
o Support production operations, monitoring, deployments, and operational readiness.
o Participate in on-call rotations and collaborate with development, infrastructure, and operations teams to improve platform reliability.
Required Skills:
o Experience leading Major Incident Management (P1/P2) in a 24x7 enterprise environment.
o Strong background in Site Reliability Engineering (SRE) or L3 Production Support.
o Experience with ServiceNow, ITIL, and enterprise monitoring tools such as Splunk, Grafana, Dynatrace, or similar.
o Hands-on troubleshooting experience with Linux/Unix, cloud platforms, databases, and distributed applications.
o Scripting experience with Python, Bash, or PowerShell is preferred.
o Excellent communication skills with the ability to coordinate technical and business stakeholders during critical incidents.
Skills
SAFe
Similar jobs
Systems Support Engineer
UnitedHealth Group · MESA, United States
17 minutes ago$72.8k - $130k/yrSystems Support Engineer
UnitedHealth Group · OVERLAND PARK, United States
17 minutes ago$72.8k - $130k/yrExecutive IT Support Lead
AVI FOODSYSTEMS INC. · Sharon, United States
1 hour ago$75k - $90k/yrExecutive IT Support Lead
AVI FOODSYSTEMS INC. · Akron, United States
1 hour ago$75k - $90k/yrIT Help Desk / Litigation Support Specialist
Ledgent Technology · Dallas, United States
1 hour ago$70k - $95k/yrIT HelpDesk Technician with Security Clearance
Titan Technologies, LLC · Derwood, United States
2 hours ago