Cloud & Compute SRE Lead (Governance & Operations Manager)
Quick Overview
Job Description
Cloud & Compute SRE Lead (Governance & Operations Manager)
About the Role
We are seeking an experienced Cloud & Compute SRE Lead to drive operational governance, reliability, and cloud infrastructure excellence across enterprise-scale environments. This leadership role is ideal for professionals who have extensive experience managing multi-cloud infrastructure, leading Site Reliability Engineering (SRE) teams, implementing automation strategies, and ensuring operational resilience.
The successful candidate will oversee Business-As-Usual (BAU) operations while driving continuous improvement through Infrastructure as Code (IaC), observability, automation, FinOps, and operational governance.
Key Responsibilities
Cloud Operations & Governance
- Define, implement, and govern SLA, SLO, and SLI frameworks across enterprise cloud platforms.
- Lead cloud financial management (FinOps), including cost optimization and resource utilization.
- Ensure compliance with cloud security standards, governance policies, and regulatory requirements.
- Manage vendor relationships and operational tooling lifecycle.
Site Reliability & Incident Management
- Act as Incident Commander during critical production incidents.
- Lead Post Incident Reviews (PIRs) and drive corrective actions.
- Own Disaster Recovery (DR) strategy, business continuity planning, and failover testing.
- Perform capacity planning and infrastructure forecasting.
Automation & Engineering Excellence
- Drive automation initiatives to eliminate manual operational tasks.
- Implement Infrastructure as Code (Terraform/OpenTofu/Ansible).
- Build governance guardrails using policy-as-code and automated compliance.
- Standardize enterprise observability using monitoring, logging, and tracing platforms.
Team Leadership
- Lead daily operations, sprint planning, standups, retrospectives, and Kanban execution.
- Manage 24x7 on-call rotations and follow-the-sun support models.
- Mentor engineers and foster a culture of reliability, automation, and continuous improvement.
Required Qualifications
- 7+ years of Infrastructure, Cloud Engineering, DevOps, or Site Reliability Engineering experience.
- 3+ years leading Infrastructure, Cloud Operations, or SRE teams.
- Strong experience managing enterprise-scale AWS, Azure, or Google Cloud environments.
- Experience with enterprise operational governance and cloud service management.
- Cloud certifications (AWS, Azure, or Google Cloud Platform) preferred.
Skills
Similar jobs
Site Reliability Engineer - W2 only
Avacend, Inc. · United States
10 minutes agoLead AWS Cloud Platform Engineer
nTech Solutions · Reston, United States
11 minutes agoStaff SAP Basis Platform Engineer
Rivian · Atlanta, United States
11 minutes ago$114.1k - $142.6k/yrSenior Devops Engineer/Lead
Infinite Computer Solutions (ICS) · Dallas, United States
11 minutes ago$80k/yrE01-L03 Cloud DevOps Engineer II with Security Clearance
EXPANSIA · Hanscom AFB, United States
1 hour ago$119k/yrNetwork Automation Engineer
Robert Half · Philadelphia, United States
1 hour ago