Haystack
← Back to Jobs
Technology
IN

Lead Site Reliability Engineer (SRE)

INGENworksUnited States🇺🇸United StatesPosted 25 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
United States
Posted
Yesterday
DockerShellAWSSplunkAnsibleAzureDatadogGoogle CloudGrafanaKubernetesPrometheusPythonTerraform

Job Description

🚀 HIRING: Journey-Centric Lead Site Reliability Engineer (SRE)

📍 Location: Hybrid (Occasional onsite visit required once or twice per month.)
💼 Duration: Long-Term Contract
🏦 Domain: Banking & Financial Services (BFS)
🎯 Experience: Lead Level

We are looking for an experienced  Lead SRE to drive end-to-end reliability, observability, automation, and operational excellence across critical customer and business journeys within the Banking & financial services domain.

🔹 Key Responsibilities

• Own reliability across critical end-to-end customer/business journeys
• Define and implement SLIs, SLOs, SLAs, and Error Budgets
• Lead observability and monitoring initiatives across applications, APIs, infrastructure, and cloud platforms
• Drive incident management, RCA, problem management, and continuous improvement
• Design and implement self-healing and automated remediation capabilities
• Leverage AI/AIOps for anomaly detection, predictive monitoring, event correlation, and intelligent operations
• Improve system availability, performance, scalability, and resilience
• Lead reliability engineering and resilience/chaos testing initiatives
• Partner with Application, Cloud, DevOps, Infrastructure, Security, and Product teams
• Establish SRE standards, best practices, dashboards, and reliability metrics

🔹 Required Skills

✅ Strong hands-on SRE / DevOps / Production Engineering experience
✅ Strong Banking & Financial Services (BFS) domain experience
✅ End-to-end journey monitoring and reliability engineering
✅ Experience with Dynatrace / Splunk / Datadog / Grafana / Prometheus
✅ Strong AWS / Azure / Google Cloud Platform experience
✅ Kubernetes & Docker
✅ Terraform / Ansible / Infrastructure as Code
✅ Python / Shell scripting for automation
✅ CI/CD and DevOps practices
✅ Strong understanding of SLO, SLI, SLA, MTTR, MTBF & Error Budgets
✅ Experience with AIOps, AI-driven operations, automation, and self-healing
✅ Strong leadership, communication, and stakeholder-management skills

🌟 Ideal Candidate:
A hands-on SRE leader who can connect technology reliability to real business/customer journeys and drive measurable improvements in availability, performance, resilience, and customer experience.

 

Similar jobs