Why This Role Stands Out
This role offers a fantastic opportunity to drive critical incident response and implement lasting improvements within a reputable cloud solutions company, with the flexibility of a hybrid work model. You'll thrive if you excel at decisive leadership under pressure and can effectively align diverse teams towards swift service restoration. Apply now to make a significant impact and advance your career in incident management.
Quick Overview
Job Description
We are looking for an experienced Major Incident Manager to lead the response to critical production incidents across our enterprise technology landscape. You will own the incident lifecycle end to end: command and coordination during the bridge, executive communication, root cause analysis, and the improvements that prevent repeat outages. The ideal candidate stays calm under pressure, makes sound decisions quickly, and can align technical teams and senior stakeholders around one goal: fast, safe service restoration.
Key Responsibilities
- Lead and coordinate major incident bridges, acting as incident commander from detection to resolution.
- Assess incident severity and business impact, and set priority in line with ITIL-aligned processes.
- Drive escalation and decision-making during high-pressure situations, engaging the right technical and vendor teams.
- Give timely, clear updates to executives and business stakeholders throughout the incident.
- Use observability tools (Splunk, Dynatrace, AppDynamics, etc.) to support triage, diagnosis, and validation of recovery.
- Lead post-incident reviews and root cause analysis (RCA), and track corrective and preventive actions to closure.
- Work with Problem, Change, and Release Management to reduce repeat incidents and change-related failures.
- Apply SRE principles (SLAs/SLOs, error budgets, toil reduction, automation) to improve reliability and operational resilience.
- Maintain incident playbooks, runbooks, and technical documentation, and produce regular reporting on incident trends and KPIs (MTTR, MTTD, recurrence).
- Support audit, compliance, and business continuity/disaster recovery requirements.
Required Skills
- Major incident management and incident command leadership
- Enterprise production support and operations management
- Critical incident response and service restoration
- Observability platforms (Splunk, Dynatrace, AppDynamics, etc.)
- Executive and stakeholder communication
- Incident severity assessment and business impact analysis
- ITIL incident management processes and governance
- Root cause analysis and post-incident reviews
- Cross-functional technical team coordination
- SRE concepts, SLA/SLO awareness, automation and toil reduction
Preferred Skills
- Problem Management and Change Management
- Monitoring and alerting tools
- Cloud infrastructure (AWS, Azure, Google Cloud Platform)
- Application support and middleware technologies
- Audit and compliance management
- Business continuity and disaster recovery
- Process improvement and operational excellence
- Technical documentation and reporting
Similar jobs
- IN
Senior Network Engineer
NewInfojini
New York, NY🇺🇸On-site17 hours agoFiberTCP/IPAnsible+8Technology - AG
SAP BASIS Consultant
NewASCII Group LLC
New York, NY🇺🇸On-site17 hours agoTechnology - MM
IT Consultant
NewMitchell Martin, Inc.
Rosemont, IL🇺🇸$44 - $54/hrHybrid17 hours agoShellAWSMLOps+7Technology - TS
Workday Finance Functional Consultant
NewTechRAQ Solutions Inc
Houston, TX🇺🇸Hybrid17 hours agoAccounts PayableAccounts ReceivableFinancial Reporting+2Technology - MM
IT Consultant - Cloud Specialist
NewMitchell Martin, Inc.
Rosemont, IL🇺🇸$47 - $57/hrHybrid17 hours agoAgileTechnology - MC
SAP GTS Consultant
NewMcKinsol Consulting Inc
United States🇺🇸Hybrid17 hours agoTechnology