Haystack
← Back to Jobs
Technology

SRE Architect

Info Way SolutionsSeattle, WA🇺🇸United StatesPosted 11 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

SRE Architect (No Visa Restriction )

Client: Cognizant
Location: Seattle, WA
Experience: 10+ Years

Job Description

We are seeking an experienced SRE Architect to lead enterprise-wide reliability engineering, observability, and operational excellence initiatives. The ideal candidate must have proven experience driving SRE transformation, defining reliability strategies, establishing SLO governance, and leading reliability engineering adoption across large-scale enterprise environments.

Key Responsibilities & Required Experience

  • Proven experience defining and implementing SLI, SLO, SLA, Error Budget, and Reliability Governance frameworks.
  • Strong expertise with Dynatrace and enterprise observability platforms, including APM, Distributed Tracing, RUM, Synthetic Monitoring, and OpenTelemetry.
  • Extensive experience with Kubernetes, Docker, OpenShift, and cloud-native architectures.
  • Design and implement High Availability (HA), Disaster Recovery (DR), Resilience, and Business Continuity strategies.
  • Lead Incident Management, Problem Management, Root Cause Analysis (RCA), and Continuous Service Improvement initiatives.
  • Drive toil reduction, self-healing, auto-remediation, and operational automation across enterprise platforms.
  • Strong understanding of Capacity Planning, Performance Engineering, and Chaos Engineering practices.
  • Establish enterprise-wide observability strategy, standards, governance, monitoring frameworks, and reliability engineering practices.
  • Experience with AIOps, predictive analytics, event correlation, intelligent alert management, and automated incident response.
  • Strong knowledge of CI/CD, DevSecOps, GitOps, and Infrastructure as Code (IaC) practices.
  • Lead Production Readiness Reviews (PRR), Operational Readiness Reviews (ORR), and enterprise reliability assessments.
  • Partner with engineering, architecture, operations, and business leadership teams to improve reliability and operational maturity.
  • Lead SRE transformation programs and drive adoption of reliability engineering principles across large-scale enterprise environments.
  • Define and track reliability KPIs, operational health metrics, SLO compliance, error budgets, and service maturity.
  • Establish frameworks for continuous improvement, operational excellence, automation, and reliability culture.

Most Important Requirement

The SRE Architect must have proven enterprise-level leadership experience, not simply experience implementing monitoring tools.

The candidate should have demonstrated success in:

  • Leading enterprise SRE transformation
  • Defining and executing enterprise observability strategy
  • Establishing and governing SLO/SLI, SLA, and Error Budget frameworks
  • Driving reliability engineering adoption across engineering organizations
  • Leading operational excellence and continuous improvement initiatives
  • Reducing operational toil through automation, self-healing, and auto-remediation
  • Influencing engineering and business leadership toward reliability-first practices
  • Building scalable SRE governance, standards, and operating models

Preferred Certifications

Relevant certifications in Azure, Kubernetes, Dynatrace, SRE, DevOps, or Cloud Architecture are preferred.

Skills

Docker
Azure
Kubernetes

Similar jobs