Why This Role Stands Out
You'll have the opportunity to lead and grow a talented SRE team, shaping the reliability strategy for customer-facing platforms and driving automation in a fully remote setting. This role is ideal for a seasoned SRE leader passionate about building blameless, data-driven cultures and eager to make a significant impact on a company's technical foundation. Embrace this chance to advance your career and contribute to a dynamic technology environment.
Quick Overview
Job Description
Location : 100 % Remote
Duration : 3 months Contract to Hire
Need only on 1099 / W2
Site Reliability Engineering Manager
SRE Manager to lead a team of reliability engineers responsible for the uptime, performance, and efficiency of the customer-facing platforms. You ll set SLOs and error budgets, build great incident and change practices, and coach engineers to automate everything that can be automated.
Responsibilities
- Lead & grow the team: Hire, coach, and develop SREs; set goals and establish a blameless, data-driven culture.
- Own reliability strategy: Define and socialize SLOs/SLIs and error budgets with product/engineering; enforce guardrails and tradeoffs.
- Operate the platform: Oversee availability, latency, capacity planning, and change management across [AWS/Azure/Google Cloud Platform] and Kubernetes.
- Incident management: Run on-call and escalation programs (SEV1/2), coordinate response, and ensure high-quality, blameless postmortems with clear follow-ups.
- Observability: Standardize logs/metrics/traces and dashboards; reduce alert noise; drive adoption of APM/monitoring tools ([Datadog/Dynatrace/PrometheGrafana/New Relic]).
- Automation & resilience: Champion infra-as-code, CI/CD, chaos/game days, load testing, and toil reduction.
- Security & compliance partnership: Work with Security, Compliance, and Finance on least-privilege, secrets management, cost efficiency, and audit readiness.
- Stakeholder alignment: Partner with Product, App Eng, Data, and Support to prioritize reliability work and communicate risk/status to leadership.
Requirements
- in software/platform/reliability engineering, including 2 4 years leading SRE/DevOps/Platform teams.
- Proven experience operating large-scale services on [AWS/Azure/Google Cloud Platform] with Kubernetes and containers.
- Strong fundamentals in Linux, networking, and distributed systems.
- Hands-on with IaC (Terraform/CloudFormation/Bicep), CI/CD (GitHub Actions/CircleCI/Azure DevOps), and one scripting language (Python/Go/Bash).
- Deep understanding of observability (metrics, logs, traces) and alerting best practices.
- Track record running on-call programs and driving measurable reliability improvements.
- Excellent communication and stakeholder management; comfortable presenting trade-offs and data to executives.
Similar jobs
- AI
GenAI / Agent Platform Engineer
NewARK Infotech Spectrum
Durham, NC🇺🇸HybridYesterdayDockerMicroservicesAWS+9Technology - CO
Principal Network & Platform Engineer / Architect - Virtualized Network Services
NewConglomerateIT
Dallas, TX🇺🇸HybridYesterdayOpenStackKubernetesRedisTechnology - TG
Site Reliability Engineer
NewTalent Groups
Minneapolis, MN🇺🇸HybridYesterdayAWSELKLogstash+3Technology - ZS
Lead ForgeRock Platform Engineer - Remote
NewZodiac Solutions Inc.
United States🇺🇸RemoteYesterdayMicroservicesSpringSpring Boot+14Technology - E-
AWS Cloud Platform Engineer / Senior AWS Infrastructure Engineer
NewE-Solutions, Inc.
Woodbridge Township, NJ🇺🇸HybridYesterdayAWSSSOBash+4Technology - TR
Role: Sr. Business Analyst with Azure DevOps Experince, Houston, TX Hybrid - Need Locals
NewTech Rakers
Houston, TX🇺🇸HybridYesterdayScrumAgileAzureTechnology