Haystack
← Back to Jobs
Remote
Technology
KA

Requirement for SRE Manager --- Remote

KairosUnited States🇺🇸United StatesPosted Sep 23, 2026

Why This Role Stands Out

You'll have the opportunity to lead and grow a talented SRE team, shaping the reliability strategy for customer-facing platforms and driving automation in a fully remote setting. This role is ideal for a seasoned SRE leader passionate about building blameless, data-driven cultures and eager to make a significant impact on a company's technical foundation. Embrace this chance to advance your career and contribute to a dynamic technology environment.

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
Yesterday
AWSNew RelicAzureBashCircleCICloudFormationDatadogGitHub ActionsGoogle CloudKubernetesPythonStakeholder ManagementTerraform

Job Description

Location : 100 % Remote

Duration : 3 months Contract to Hire

Need only on 1099 / W2

Site Reliability Engineering Manager

SRE Manager to lead a team of reliability engineers responsible for the uptime, performance, and efficiency of the customer-facing platforms. You ll set SLOs and error budgets, build great incident and change practices, and coach engineers to automate everything that can be automated.

Responsibilities

  • Lead & grow the team: Hire, coach, and develop SREs; set goals and establish a blameless, data-driven culture.
  • Own reliability strategy: Define and socialize SLOs/SLIs and error budgets with product/engineering; enforce guardrails and tradeoffs.
  • Operate the platform: Oversee availability, latency, capacity planning, and change management across [AWS/Azure/Google Cloud Platform] and Kubernetes.
  • Incident management: Run on-call and escalation programs (SEV1/2), coordinate response, and ensure high-quality, blameless postmortems with clear follow-ups.
  • Observability: Standardize logs/metrics/traces and dashboards; reduce alert noise; drive adoption of APM/monitoring tools ([Datadog/Dynatrace/PrometheGrafana/New Relic]).
  • Automation & resilience: Champion infra-as-code, CI/CD, chaos/game days, load testing, and toil reduction.
  • Security & compliance partnership: Work with Security, Compliance, and Finance on least-privilege, secrets management, cost efficiency, and audit readiness.
  • Stakeholder alignment: Partner with Product, App Eng, Data, and Support to prioritize reliability work and communicate risk/status to leadership.

Requirements

  • in software/platform/reliability engineering, including 2 4 years leading SRE/DevOps/Platform teams.
  • Proven experience operating large-scale services on [AWS/Azure/Google Cloud Platform] with Kubernetes and containers.
  • Strong fundamentals in Linux, networking, and distributed systems.
  • Hands-on with IaC (Terraform/CloudFormation/Bicep), CI/CD (GitHub Actions/CircleCI/Azure DevOps), and one scripting language (Python/Go/Bash).
  • Deep understanding of observability (metrics, logs, traces) and alerting best practices.
  • Track record running on-call programs and driving measurable reliability improvements.
  • Excellent communication and stakeholder management; comfortable presenting trade-offs and data to executives.

Similar jobs