Why This Role Stands Out
This leadership role offers a fantastic opportunity to drive significant improvements in cloud-native application reliability and scalability, making a tangible impact on a reputable company. If you are a seasoned SRE professional who thrives on technical challenges and leading innovation, you will find this position incredibly rewarding. Apply today to leverage your expertise and shape the future of their SRE organization!
Quick Overview
Job Description
Senior / Lead Site Reliability Engineer, Principal Engineer
Long term contract
Phoenix, AZ (Hybrid-3 days onsite)
Direct client- Immediate client interview
Role Overview
We are looking for a highly experienced Senior / Lead Site Reliability Engineer to serve as a technical leader and go-to engineer for the SRE organization. This role will focus on improving the reliability, observability, scalability, deployment safety, and operational readiness of cloud-native applications.
The ideal candidate brings a strong SRE mindset, deep production engineering experience, and the ability to identify reliability gaps, challenge existing approaches, and drive improvements across engineering teams.
Role is local to Phx, or someone willing to relocate.
Required Qualifications
Strong Site Reliability Engineering experience supporting highly available, production-scale systems.
Strong hands-on Google Cloud Platform (Google Cloud Platform) experience.
Deep understanding and practical application of SLIs, SLOs, error budgets, operability, and reliability engineering principles.
Strong experience with observability and instrumentation, including metrics, logging, tracing, alerting, and production diagnostics.
Experience with production troubleshooting, incident response, root-cause analysis, and operational readiness.
Experience with Terraform or another Infrastructure as Code technology.
Experience developing and improving CI/CD pipelines and deployment practices.
Proficiency with Python, Bash, or another scripting/automation language.
Ability to identify systemic reliability issues and drive engineering solutions rather than primarily responding to operational incidents.
Strong technical leadership, collaboration, and communication skills across application, platform, and engineering teams.
Preferred Qualifications
Experience supporting Kubernetes and GKE workloads in production.
Experience with GitHub Actions.
Experience with Google Cloud Platform Cloud Monitoring, OpenTelemetry, Prometheus, or Grafana.
Experience with Istio or another service mesh.
Experience with Apigee or another API gateway.
Experience implementing canary, blue/green, or progressive delivery strategies.
Experience improving engineering automation and automated testing practices.
This is a hands-on technical leadership role with the opportunity to become a key technical authority for SRE practices across the organization.
Similar jobs
- PS
DevOps Cloud Engineer with Security Clearance
Power3 Solutions
Annapolis Junction, MD🇺🇸$140k - $235k/yrHybrid6 weeks agoDockerRubyAWS+8Technology - IS
Principal Engineer – DevSecOps & Release Engineering
Innova Solutions, Inc
Irving, TX🇺🇸$80 - $90/hrHybrid1 week agoDjangoFastAPIFlask+8Technology - GD
DevOps Automation Engineer
GDH
United States🇺🇸$58 - $61/hrRemote1 week agoAWSAnsibleBash+7Technology - GO
Cyber / DevOps Systems Engineer (Ft. Bragg)
NewGovcio LLC
Fort Bragg, NC🇺🇸$120k - $130k/yrOn-siteYesterdayEncryptionSplunkAnsible+4Technology - SW
DevSecOps Engineer
NewSTS Worldwide Inc.
Washington, DC🇺🇸Hybrid9 hours agoEngineering - BO
Configuration Management Engineer with Security Clearance
NewBoeing
Oklahoma City, OK🇺🇸$79.0k - $107.0k/yrOn-siteYesterdayJiraEngineering