Haystack
← Back to Jobs
Other
SC

Major Incident Manager

SN Cloud SolutionsPhoenix, AZ🇺🇸United StatesPosted Sep 29, 2026

Why This Role Stands Out

This role offers a fantastic opportunity to drive critical incident response and implement lasting improvements within a reputable cloud solutions company, with the flexibility of a hybrid work model. You'll thrive if you excel at decisive leadership under pressure and can effectively align diverse teams towards swift service restoration. Apply now to make a significant impact and advance your career in incident management.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Phoenix, AZ, United States
Posted
17 hours ago
AWSSplunkAzureComplianceGoogle CloudProcess ImprovementRoot Cause AnalysisTriage

Job Description

We are looking for an experienced Major Incident Manager to lead the response to critical production incidents across our enterprise technology landscape. You will own the incident lifecycle end to end: command and coordination during the bridge, executive communication, root cause analysis, and the improvements that prevent repeat outages. The ideal candidate stays calm under pressure, makes sound decisions quickly, and can align technical teams and senior stakeholders around one goal: fast, safe service restoration.

Key Responsibilities

  • Lead and coordinate major incident bridges, acting as incident commander from detection to resolution.
  • Assess incident severity and business impact, and set priority in line with ITIL-aligned processes.
  • Drive escalation and decision-making during high-pressure situations, engaging the right technical and vendor teams.
  • Give timely, clear updates to executives and business stakeholders throughout the incident.
  • Use observability tools (Splunk, Dynatrace, AppDynamics, etc.) to support triage, diagnosis, and validation of recovery.
  • Lead post-incident reviews and root cause analysis (RCA), and track corrective and preventive actions to closure.
  • Work with Problem, Change, and Release Management to reduce repeat incidents and change-related failures.
  • Apply SRE principles (SLAs/SLOs, error budgets, toil reduction, automation) to improve reliability and operational resilience.
  • Maintain incident playbooks, runbooks, and technical documentation, and produce regular reporting on incident trends and KPIs (MTTR, MTTD, recurrence).
  • Support audit, compliance, and business continuity/disaster recovery requirements.

Required Skills

  • Major incident management and incident command leadership
  • Enterprise production support and operations management
  • Critical incident response and service restoration
  • Observability platforms (Splunk, Dynatrace, AppDynamics, etc.)
  • Executive and stakeholder communication
  • Incident severity assessment and business impact analysis
  • ITIL incident management processes and governance
  • Root cause analysis and post-incident reviews
  • Cross-functional technical team coordination
  • SRE concepts, SLA/SLO awareness, automation and toil reduction

Preferred Skills

  • Problem Management and Change Management
  • Monitoring and alerting tools
  • Cloud infrastructure (AWS, Azure, Google Cloud Platform)
  • Application support and middleware technologies
  • Audit and compliance management
  • Business continuity and disaster recovery
  • Process improvement and operational excellence
  • Technical documentation and reporting

Similar jobs