Haystack
← Back to Jobs
Technology

BizOps SRE

Hire Tech ServicesSt. Louis, MO🇺🇸United StatesPosted 11 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

BizOps SRE – Production Engineering

Location: St. Louis, MO – Hybrid (3 days onsite per week)
Job Type: Long-Term Contract

Job Overview

We are seeking an experienced BizOps SRE / Production Engineer to support critical production services and drive reliability, automation, and operational excellence across a global environment.

The ideal candidate will have strong hands-on experience with SRE, DevOps, production operations, cloud infrastructure, Kubernetes, Terraform, observability, incident management, and automation. This role will focus on improving service stability, reducing MTTR and operational toil, strengthening production readiness, and partnering closely with Engineering, SRE, Operations, and Product teams.

Key Responsibilities

  • Own production operations, service stability, availability, and continuity for critical services.
  • Lead Major Incidents / Technical Response Team (TRT) activities and drive timely resolution and reduced MTTR.
  • Provide clear and timely communication to stakeholders, customers, and leadership during production incidents.
  • Drive operational readiness for application releases, infrastructure migrations, and peak business events.
  • Implement and improve monitoring, observability, alerting, and service health dashboards.
  • Lead Root Cause Analysis (RCA) and problem management activities.
  • Identify reliability risks and implement proactive solutions to improve system resilience.
  • Support capacity planning, performance improvement, and resilience initiatives.
  • Automate repetitive operational activities and reduce manual operational toil.
  • Develop and maintain reusable runbooks, SOPs, operational procedures, and readiness standards.
  • Partner with Engineering, Infrastructure SRE, Operations, Product, and other technical teams.
  • Convert production insights and incident learnings into engineering improvements and roadmap initiatives.
  • Participate in global production support and operational activities as required.

Required Skills

  • Strong experience in SRE, DevOps, Production Engineering, or Site Reliability Engineering.
  • Hands-on experience supporting production environments.
  • Strong knowledge of Linux administration and troubleshooting.
  • Experience with at least one major cloud platform: AWS, Azure, or Google Cloud Platform.
  • Hands-on experience with Kubernetes.
  • Experience with Terraform and Infrastructure as Code.
  • Strong scripting/programming experience with one or more of:
    • Python
    • Go
    • Java
    • Bash
  • Experience with observability and monitoring tools such as:
    • Splunk
    • Dynatrace
    • Prometheus
    • Grafana
    • Datadog
  • Strong experience with Incident Management, Major Incident Management, RCA, Problem Management, and Change/Release Management.
  • Experience with CI/CD pipelines and automation.
  • Strong troubleshooting and production support skills.
  • Excellent communication and stakeholder management skills.

Preferred Qualifications

  • Experience supporting 24x7/global production environments.
  • Experience working with distributed/global teams.
  • Experience in financial services, banking, payments, or other high-availability environments.
  • Experience with production readiness reviews and release/migration readiness.
  • Experience developing operational runbooks, SOPs, and reliability standards.
  • Demonstrated experience reducing incidents, MTTR, or operational toil through automation and process improvements.

Ideal Candidate

The ideal candidate is a hands-on SRE/Production Engineer who can operate effectively in a high-availability production environment, lead critical incidents, troubleshoot complex infrastructure and application issues, automate repetitive processes, and collaborate effectively with engineering and business stakeholders.

Skills

AWS
Splunk
Azure
Bash
Datadog
Google Cloud
Grafana
Java
Kubernetes
Prometheus
Python
Stakeholder Management
Terraform

Similar jobs