Quick Overview
Salary
$173k - $230k/yr
Seniority
Leader
Work mode
Hybrid
Location
Scottsdale, AZ, United States
Posted
11 hours ago
Phoenix
Job Description
At Early Warning, we've powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle , Paze , and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.
Early Warning follows a hybrid work model to allow for a more collaborative working environment.
Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.
Role Summary
The Director, Site Reliability Engineering leads the SRE function for an assigned product, platform, pillar or business domain and is accountable for the reliability, scalability, performance and operability of its business-critical production services. The Director develops engineering talent, establishes domain reliability strategy and priorities, and partners with Product, Software Engineering, Platform, Infrastructure, Security and Risk teams relevant to the domain.
The role translates business and product priorities into measurable reliability outcomes and ensures teams have the engineering practices, capabilities and operating mechanisms needed to achieve them.
Core SRE Responsibilities
1. Reliability Objectives and Measurement
Establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) aligned with customer, product and business outcomes.
Use quantitative production data to measure performance and partner with domain teams to set reliability targets appropriate to business criticality, architecture, customer expectations and cost.
2. Reliability Risk and Error Budgets
Use SLO performance, error budgets, failure data and production evidence to identify and prioritize reliability risk.
Ensure persistent risks have accountable engineering plans, appropriate escalation and sustained follow-through.
3. Software Engineering and Automation
Champion reusable software, tooling and automation that improve reliability, scalability and operability while reducing repetitive operational work.
Ensure SRE-developed solutions follow sound engineering practices, including source control, code review, testing, maintainability and secure development.
The Director is expected to remain hands-on and maintain sufficient technical depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.
4. Observability Engineering
Establish expectations for metrics, logs, traces, dashboards and telemetry sufficient to understand service behavior, customer impact and systemic risk.
Drive actionable alerting and observability that support rapid diagnosis, performance analysis, capacity planning and continuous improvement.
5. Incident Management and Service Restoration
Provide domain leadership and accountability for disciplined incident response, service restoration and stakeholder communication.
Improve response through automation, runbooks, training, exercises and analysis of recurring patterns.
6. Learning From Failure
Promote blameless post-incident review focused on systemic learning, prevention and measurable follow-through.
Use patterns across incidents and near misses to drive architectural, operational and domain-level improvements.
7. Production Readiness, Resilience and Capacity
Ensure critical services meet appropriate production-readiness, resilience, recovery, capacity and scalability expectations.
Partner with engineering teams to address systemic failure modes, dependency risks, capacity constraints and recovery gaps before they affect customers.
8. Operational Toil and Sustainable Engineering
Measure and reduce repetitive, manual and low-value operational work through engineering, automation and simplification.
Ensure operational responsibilities inform engineering priorities without becoming the primary definition of the SRE role or creating unsustainable team load.
9. Security, Risk and Compliance
Partner with Security, Risk, Compliance and engineering teams relevant to the domain to meet EWS control, resilience and regulatory obligations.
Ensure material reliability risks and control gaps are visible and addressed within the domain or escalated when they exceed its authority.
The Director operates within an assigned domain and delivers outcomes through its SRE teams. Impact is demonstrated by building strong teams and leaders, setting direction, resolving barriers within the domain and with direct dependencies, and creating durable engineering mechanisms. The Director leads through others while remaining sufficiently hands-on to contribute directly when appropriate. Success is measured primarily by the reliability outcomes and capabilities of the teams the Director leads.
Leadership Responsibilities
Own SRE strategy, execution and measurable reliability outcomes for the assigned domain.
Build, develop and retain high-performing SRE teams with clear accountability, career development and succession.
Translate business and product priorities into reliability investments and an executable multi-quarter roadmap.
Set priorities and make evidence-based tradeoffs across reliability, delivery, operational risk, capacity and business needs.
Develop capable managers and technical leaders, ensuring decisions are made at the appropriate level.
Partner with engineering, product, infrastructure, security and risk teams relevant to the domain to embed reliability into engineering decisions.
Make reliability risk, SLO performance, operational load and improvement progress visible through effective metrics and operating reviews.
Provide accountable leadership during significant incidents and ensure systemic corrective actions are completed.
Required Qualifications
Typically 12 + years of relevant software engineering, site reliability engineering, production engineering, platform engineering or closely related experience, including significant technical leadership.
5+ years of people leadership experience, with demonstrated success leading engineering teams and developing managers and/or senior technical leaders.
Experience operating highly available, business-critical distributed systems and leading reliability improvement across a major product, platform, pillar or business domain.
Strong understanding of SRE practices, including SLOs/SLIs, error budgets, observability, incident management, automation, capacity, resilience and production readiness.
Demonstrated hands-on technical capability and sufficient depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.
Ability to communicate technical risk, tradeoffs and investment needs clearly to engineering, product and senior business stakeholders.
Demonstrated ability to build inclusive, accountable and high-performing engineering teams.
Preferred Qualifications
Experience in payments, financial services or another highly regulated, high-availability environment.
Experience leading SRE or production engineering across multiple teams or a complex product or platform ecosystem.
Experience with cloud platforms, distributed systems, modern observability, infrastructure automation and software delivery at scale.
Experience establishing reliability metrics, governance and operating reviews across teams within a defined domain.
The base pay scale for this position in:
Phoenix, AZ/ Chicago, IL in USD per year is: $173,000 - $230,000.
San Francisco, CA in USD per year is: $207,000 - $276,000.
Additionally, candidates are eligible for a discretionary incentive plan and benefits.
Some of the Ways We Prioritize Your Health and Happiness
And SO much more! We continue to enhance our program, so be sure to check our Benefits page here for the latest. Our team can share more during the interview process!
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Early Warning Services, LLC ("Early Warning") considers for employment, hires, retains and promotes qualified candidates on the basis of ability, potential, and valid qualifications without regard to race, religious creed, religion, color, sex, sexual orientation, genetic information, gender, gender identity, gender expression, age, national origin, ancestry, citizenship, protected veteran or disability status or any factor prohibited by law, and as such affirms in policy and practice to support and promote equal employment opportunity and affirmative action, in accordance with all applicable federal, state, and municipal laws. The company also prohibits discrimination on other bases such as medical condition, marital status or any other factor that is irrelevant to the performance of our employees.
Early Warning follows a hybrid work model to allow for a more collaborative working environment.
Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.
Role Summary
The Director, Site Reliability Engineering leads the SRE function for an assigned product, platform, pillar or business domain and is accountable for the reliability, scalability, performance and operability of its business-critical production services. The Director develops engineering talent, establishes domain reliability strategy and priorities, and partners with Product, Software Engineering, Platform, Infrastructure, Security and Risk teams relevant to the domain.
The role translates business and product priorities into measurable reliability outcomes and ensures teams have the engineering practices, capabilities and operating mechanisms needed to achieve them.
Core SRE Responsibilities
1. Reliability Objectives and Measurement
Establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) aligned with customer, product and business outcomes.
Use quantitative production data to measure performance and partner with domain teams to set reliability targets appropriate to business criticality, architecture, customer expectations and cost.
2. Reliability Risk and Error Budgets
Use SLO performance, error budgets, failure data and production evidence to identify and prioritize reliability risk.
Ensure persistent risks have accountable engineering plans, appropriate escalation and sustained follow-through.
3. Software Engineering and Automation
Champion reusable software, tooling and automation that improve reliability, scalability and operability while reducing repetitive operational work.
Ensure SRE-developed solutions follow sound engineering practices, including source control, code review, testing, maintainability and secure development.
The Director is expected to remain hands-on and maintain sufficient technical depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.
4. Observability Engineering
Establish expectations for metrics, logs, traces, dashboards and telemetry sufficient to understand service behavior, customer impact and systemic risk.
Drive actionable alerting and observability that support rapid diagnosis, performance analysis, capacity planning and continuous improvement.
5. Incident Management and Service Restoration
Provide domain leadership and accountability for disciplined incident response, service restoration and stakeholder communication.
Improve response through automation, runbooks, training, exercises and analysis of recurring patterns.
6. Learning From Failure
Promote blameless post-incident review focused on systemic learning, prevention and measurable follow-through.
Use patterns across incidents and near misses to drive architectural, operational and domain-level improvements.
7. Production Readiness, Resilience and Capacity
Ensure critical services meet appropriate production-readiness, resilience, recovery, capacity and scalability expectations.
Partner with engineering teams to address systemic failure modes, dependency risks, capacity constraints and recovery gaps before they affect customers.
8. Operational Toil and Sustainable Engineering
Measure and reduce repetitive, manual and low-value operational work through engineering, automation and simplification.
Ensure operational responsibilities inform engineering priorities without becoming the primary definition of the SRE role or creating unsustainable team load.
9. Security, Risk and Compliance
Partner with Security, Risk, Compliance and engineering teams relevant to the domain to meet EWS control, resilience and regulatory obligations.
Ensure material reliability risks and control gaps are visible and addressed within the domain or escalated when they exceed its authority.
The Director operates within an assigned domain and delivers outcomes through its SRE teams. Impact is demonstrated by building strong teams and leaders, setting direction, resolving barriers within the domain and with direct dependencies, and creating durable engineering mechanisms. The Director leads through others while remaining sufficiently hands-on to contribute directly when appropriate. Success is measured primarily by the reliability outcomes and capabilities of the teams the Director leads.
Leadership Responsibilities
Own SRE strategy, execution and measurable reliability outcomes for the assigned domain.
Build, develop and retain high-performing SRE teams with clear accountability, career development and succession.
Translate business and product priorities into reliability investments and an executable multi-quarter roadmap.
Set priorities and make evidence-based tradeoffs across reliability, delivery, operational risk, capacity and business needs.
Develop capable managers and technical leaders, ensuring decisions are made at the appropriate level.
Partner with engineering, product, infrastructure, security and risk teams relevant to the domain to embed reliability into engineering decisions.
Make reliability risk, SLO performance, operational load and improvement progress visible through effective metrics and operating reviews.
Provide accountable leadership during significant incidents and ensure systemic corrective actions are completed.
Required Qualifications
Typically 12 + years of relevant software engineering, site reliability engineering, production engineering, platform engineering or closely related experience, including significant technical leadership.
5+ years of people leadership experience, with demonstrated success leading engineering teams and developing managers and/or senior technical leaders.
Experience operating highly available, business-critical distributed systems and leading reliability improvement across a major product, platform, pillar or business domain.
Strong understanding of SRE practices, including SLOs/SLIs, error budgets, observability, incident management, automation, capacity, resilience and production readiness.
Demonstrated hands-on technical capability and sufficient depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.
Ability to communicate technical risk, tradeoffs and investment needs clearly to engineering, product and senior business stakeholders.
Demonstrated ability to build inclusive, accountable and high-performing engineering teams.
Preferred Qualifications
Experience in payments, financial services or another highly regulated, high-availability environment.
Experience leading SRE or production engineering across multiple teams or a complex product or platform ecosystem.
Experience with cloud platforms, distributed systems, modern observability, infrastructure automation and software delivery at scale.
Experience establishing reliability metrics, governance and operating reviews across teams within a defined domain.
The base pay scale for this position in:
Phoenix, AZ/ Chicago, IL in USD per year is: $173,000 - $230,000.
San Francisco, CA in USD per year is: $207,000 - $276,000.
Additionally, candidates are eligible for a discretionary incentive plan and benefits.
Some of the Ways We Prioritize Your Health and Happiness
- Healthcare Coverage -Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.
- 401(k) Retirement Plan -Featuring a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.
- Paid Time Off - Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.
- 12 weeks of Paid Parental Leave
- Maven Family Planning - provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.
And SO much more! We continue to enhance our program, so be sure to check our Benefits page here for the latest. Our team can share more during the interview process!
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Early Warning Services, LLC ("Early Warning") considers for employment, hires, retains and promotes qualified candidates on the basis of ability, potential, and valid qualifications without regard to race, religious creed, religion, color, sex, sexual orientation, genetic information, gender, gender identity, gender expression, age, national origin, ancestry, citizenship, protected veteran or disability status or any factor prohibited by law, and as such affirms in policy and practice to support and promote equal employment opportunity and affirmative action, in accordance with all applicable federal, state, and municipal laws. The company also prohibits discrimination on other bases such as medical condition, marital status or any other factor that is irrelevant to the performance of our employees.
Similar jobs
- BA
DevSecOps Engineer
BOOZ, ALLEN & HAMILTON, INC.
Fayetteville, NC🇺🇸$77.5k - $176k/yrOn-site2 weeks agoSAFeScrumAgileEngineering - SA
Lead DevOps Engineer
NewSAIC
MD🇺🇸$200.0k - $240k/yrRemote11 hours agoMongoDBAWSLogstash+7Technology - WO
Sr Software Development Engineer, SRE (US Federal) with Security Clearance
NewWorkday
Reston, VA🇺🇸$163.8k/yrHybridYesterdayDockerAWSSplunk+8Technology - AG
Azure Cloud Platform Engineer with Security Clearance
NewAgensys Corporation
Houston, TX🇺🇸HybridYesterdayAzureGitHub ActionsTerraformTechnology - PS
Mid-Level DevOps Software Engineer, SE2 with Security Clearance
NewPower3 Solutions
Hanover, MD🇺🇸$206k - $239k/yrHybridYesterdayMongoDBLogstashAgile+9Technology - XT
Senior Systems Engineer/Senior DevOps Engineer with Security Clearance
NewX Technologies, Inc
San Antonio, TX🇺🇸HybridYesterdayDockerAnsibleBash+5Technology - ON
Site Reliability Engineer II
NewAuto ApplyOnapsis
Dallas🇺🇸Hybrid17 hours agoOracleAWSTDD+10Technology - NE
Site Reliability Engineer
NewAuto ApplyNebius
Remote - United States🇺🇸$130k - $180k/yrRemote9 hours agoBashPythonTechnology - EL
Senior DevOps Engineer (NOAA badge required)
NewAuto ApplyElement84
Alexandria HQ (remote)🇺🇸$145k - $180k/yrRemote17 hours agoDockerDynamoDBSQL+20Technology - CM
Senior AI Platform Engineer
NewAuto ApplyCode Metal
Boston Hub🇺🇸Hybrid7 hours agoAPI GatewayMLflowSalesforce+7Technology - CL
Platform Engineer (Kubernetes), Mid-Level
NewAuto ApplyClera
San Francisco🇺🇸$180k - $210k/yrOn-site18 hours agoService MeshArgoCDGrafana+4Technology - WE
Site Reliability Engineer, Cloud Infrastructure
NewAuto ApplyWeave
Weave - Headquarters (Lehi🇺🇸Remote17 hours agoDockerGCPAnsible+11Technology