Haystack
← Back to Jobs
Technology

Cloud & Compute SRE Lead (Governance & Operations Manager)

Echo IT Solutions, Inc.Princeton, NJ🇺🇸United StatesPosted 8 Jul 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Cloud & Compute SRE Lead (Governance & Operations Manager)

About the Role

We are seeking an experienced Cloud & Compute SRE Lead to drive operational governance, reliability, and cloud infrastructure excellence across enterprise-scale environments. This leadership role is ideal for professionals who have extensive experience managing multi-cloud infrastructure, leading Site Reliability Engineering (SRE) teams, implementing automation strategies, and ensuring operational resilience.

The successful candidate will oversee Business-As-Usual (BAU) operations while driving continuous improvement through Infrastructure as Code (IaC), observability, automation, FinOps, and operational governance.


Key Responsibilities

Cloud Operations & Governance

  • Define, implement, and govern SLA, SLO, and SLI frameworks across enterprise cloud platforms.
  • Lead cloud financial management (FinOps), including cost optimization and resource utilization.
  • Ensure compliance with cloud security standards, governance policies, and regulatory requirements.
  • Manage vendor relationships and operational tooling lifecycle.

Site Reliability & Incident Management

  • Act as Incident Commander during critical production incidents.
  • Lead Post Incident Reviews (PIRs) and drive corrective actions.
  • Own Disaster Recovery (DR) strategy, business continuity planning, and failover testing.
  • Perform capacity planning and infrastructure forecasting.

Automation & Engineering Excellence

  • Drive automation initiatives to eliminate manual operational tasks.
  • Implement Infrastructure as Code (Terraform/OpenTofu/Ansible).
  • Build governance guardrails using policy-as-code and automated compliance.
  • Standardize enterprise observability using monitoring, logging, and tracing platforms.

Team Leadership

  • Lead daily operations, sprint planning, standups, retrospectives, and Kanban execution.
  • Manage 24x7 on-call rotations and follow-the-sun support models.
  • Mentor engineers and foster a culture of reliability, automation, and continuous improvement.

Required Qualifications

  • 7+ years of Infrastructure, Cloud Engineering, DevOps, or Site Reliability Engineering experience.
  • 3+ years leading Infrastructure, Cloud Operations, or SRE teams.
  • Strong experience managing enterprise-scale AWS, Azure, or Google Cloud environments.
  • Experience with enterprise operational governance and cloud service management.
  • Cloud certifications (AWS, Azure, or Google Cloud Platform) preferred.

Skills

AWS
Ansible
Azure
Google Cloud
Kanban
Terraform

Similar jobs