Haystack
← Back to Jobs
Full time
Technology
MA

Manager, Site Reliability Engineering

MastercardO Fallon, Missouri🇺🇸United StatesPosted 22 Aug 2026

Why This Role Stands Out

This Manager, Site Reliability Engineering role at Mastercard offers a fantastic opportunity to lead a team, shape SRE best practices, and drive innovation within a globally recognized company, with the added flexibility of a hybrid work model. You will thrive here if you are passionate about building highly available, scalable systems, mentoring engineers, and advancing your skills in cloud-native technologies. Apply now to contribute to cutting-edge solutions and grow your career in a collaborative and learning-focused environment.

Quick Overview

Seniority
Mid Senior
Employment type
Full Time
Work mode
Hybrid
Location
O Fallon, Missouri, United States
GCPAWSNew RelicAzureBashGrafanaKubernetesPrometheusPythonTerraform

Job Description

Mastercard seeks a Manager, Site Reliability Engineering to lead a team ensuring secure, scalable, and highly available platforms. You will design and implement SRE best practices, build automation for deployment and operations, and drive observability across complex cloud-native systems. Partner with development and security teams to improve reliability, performance, and incident response. You'll mentor engineers, champion continuous improvement, and help shape a culture of innovation, collaboration, and learning while working with cutting-edge technologies in a global environment.

Responsibilities

  • Lead and mentor an SRE team supporting mission-critical platforms
  • Define and implement SRE best practices for reliability, scalability, and security
  • Design automation for deployments, configuration, and operations
  • Establish and improve monitoring, logging, and alerting for cloud-native systems
  • Drive incident management, root-cause analysis, and post-incident reviews
  • Collaborate with software engineering and security teams to improve system design
  • Optimize performance and capacity planning across services
  • Promote continuous improvement and a learning culture within the team

Required Skills

  • Site Reliability Engineering (SRE)
  • Cloud platforms (AWS/Azure/GCP)
  • Kubernetes & container orchestration
  • Linux systems administration
  • CI/CD pipelines
  • Infrastructure as Code (Terraform/Cloud
  • Formation)
  • Monitoring & observability (Prometheus/Grafana/New Relic)
  • Incident management & on-call operations
  • Performance tuning & capacity planning
  • Scripting (Python/Bash)

Similar jobs