Quick Overview
Job Description
SRE Architect (No Visa Restriction )
Client: Cognizant
Location: Seattle, WA
Experience: 10+ Years
Job Description
We are seeking an experienced SRE Architect to lead enterprise-wide reliability engineering, observability, and operational excellence initiatives. The ideal candidate must have proven experience driving SRE transformation, defining reliability strategies, establishing SLO governance, and leading reliability engineering adoption across large-scale enterprise environments.
Key Responsibilities & Required Experience
- Proven experience defining and implementing SLI, SLO, SLA, Error Budget, and Reliability Governance frameworks.
- Strong expertise with Dynatrace and enterprise observability platforms, including APM, Distributed Tracing, RUM, Synthetic Monitoring, and OpenTelemetry.
- Extensive experience with Kubernetes, Docker, OpenShift, and cloud-native architectures.
- Design and implement High Availability (HA), Disaster Recovery (DR), Resilience, and Business Continuity strategies.
- Lead Incident Management, Problem Management, Root Cause Analysis (RCA), and Continuous Service Improvement initiatives.
- Drive toil reduction, self-healing, auto-remediation, and operational automation across enterprise platforms.
- Strong understanding of Capacity Planning, Performance Engineering, and Chaos Engineering practices.
- Establish enterprise-wide observability strategy, standards, governance, monitoring frameworks, and reliability engineering practices.
- Experience with AIOps, predictive analytics, event correlation, intelligent alert management, and automated incident response.
- Strong knowledge of CI/CD, DevSecOps, GitOps, and Infrastructure as Code (IaC) practices.
- Lead Production Readiness Reviews (PRR), Operational Readiness Reviews (ORR), and enterprise reliability assessments.
- Partner with engineering, architecture, operations, and business leadership teams to improve reliability and operational maturity.
- Lead SRE transformation programs and drive adoption of reliability engineering principles across large-scale enterprise environments.
- Define and track reliability KPIs, operational health metrics, SLO compliance, error budgets, and service maturity.
- Establish frameworks for continuous improvement, operational excellence, automation, and reliability culture.
Most Important Requirement
The SRE Architect must have proven enterprise-level leadership experience, not simply experience implementing monitoring tools.
The candidate should have demonstrated success in:
- Leading enterprise SRE transformation
- Defining and executing enterprise observability strategy
- Establishing and governing SLO/SLI, SLA, and Error Budget frameworks
- Driving reliability engineering adoption across engineering organizations
- Leading operational excellence and continuous improvement initiatives
- Reducing operational toil through automation, self-healing, and auto-remediation
- Influencing engineering and business leadership toward reliability-first practices
- Building scalable SRE governance, standards, and operating models
Preferred Certifications
Relevant certifications in Azure, Kubernetes, Dynatrace, SRE, DevOps, or Cloud Architecture are preferred.
Skills
Similar jobs
Systems Engineer
TEKsystems c/o Allegis Group · El Segundo, United States
39 minutes ago$75 - $115/hrCloud Infrastructure Site Reliability Engineer
Judge Group, Inc. · Berkeley Heights, United States
40 minutes ago$70 - $80/hrSenior Application Support Engineer (SRE)
DTCC · Coppell, United States
40 minutes agoDevOps Engineer Senior
Chenega MIOS · Reston, United States
40 minutes agoSr. DevOps Engineer
Apex Systems · Overland Park, United States
49 minutes agoSenior DevOps Engineer AWS, Kubernetes, Hybrid - 69650
PRIMUS Global Services Inc. · United States
50 minutes ago