Haystack
โ† Back to Jobs
Full time
Technology
WA

Senior Site Reliability Engineer

Weekday AIBengaluru, Karnataka๐Ÿ‡ฎ๐Ÿ‡ณIndiaPosted 19 Sept 2026

Quick Overview

Seniority
Mid Senior
Employment type
Full Time
Work mode
On Site
Location
Bengaluru, Karnataka, India
Posted
Yesterday
DockerAWSLoad BalancingPCI DSSSOC 2TCP/IPBashCapacity PlanningComplianceContinuous ImprovementDNSGrafanaHIPAAHelmKubernetesPrometheusPythonTerraform

Job Description

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿญ๐Ÿฏ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿญ๐Ÿฏ-๐Ÿฎ๐Ÿฌ ๐—Ÿ๐—ฃ๐—”)

Experience: 4+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are looking for an experiencedย Senior Site Reliability Engineer (SRE)ย to build, operate, and continuously improve highly reliable, scalable, secure, and high-performing production systems acrossย hybrid and multi-cloud environments.

The role combines cloud infrastructure, Kubernetes, automation, observability, incident management, and reliability engineering. The ideal candidate will have strong hands-on experience withย AWS, Kubernetes, Terraform, Python, Bash, and modern observability platforms, along with a strong understanding of production operations and distributed systems.

Key Responsibilities

  • Define and manageย SLIs, SLOs, SLAs, error budgets, and reliability objectivesย for critical production services.
  • Drive initiatives to improve system availability, scalability, performance, resilience, and operational efficiency.
  • Manage and support productionย Kubernetes environments, including Amazon EKS and Red Hat OpenShift.
  • Deploy and maintain containerised workloads usingย Docker, Kubernetes, and Helm.
  • Manage cloud infrastructure acrossย AWS and IBM Cloud, including hybrid-cloud environments.
  • Design and maintain reliable cloud connectivity, networking, disaster-recovery, and failover solutions.
  • Develop and maintain infrastructure usingย Terraform and Infrastructure as Code (IaC)ย practices.
  • Automate operational processes, infrastructure tasks, and troubleshooting workflows usingย Python and Bash.
  • Build and enhance observability solutions usingย Prometheus, Grafana, OpenTelemetry, Thanos, and logging platforms.
  • Monitor system health, identify performance bottlenecks, and proactively address reliability risks.
  • Participate in and leadย high-severity incident responseย and production troubleshooting.
  • Conduct root-cause analysis and lead post-incident reviews and corrective actions.
  • Develop and maintain capacity-planning and reliability-improvement strategies.
  • Implement secure, resilient, and compliant infrastructure practices across cloud environments.
  • Support disaster-recovery planning, testing, and continuous improvement.
  • Collaborate with software engineering, platform, security, and architecture teams to improve production reliability.
  • Contribute to architecture reviews, engineering standards, operational best practices, and automation initiatives.
  • Mentor engineers and promote strong SRE, DevOps, observability, and production-engineering practices.

What Makes You a Great Fit

  • 4โ€“6 years of professional experienceย in Site Reliability Engineering, DevOps, Cloud Infrastructure, or a closely related field.
  • Strong hands-on experience withย AWS and Kubernetesย in production environments.
  • Experience managingย Amazon EKS, Docker, and Helm.
  • Practical experience withย Red Hat OpenShiftย is highly desirable.
  • Strong proficiency inย Terraformย and Infrastructure as Code practices.
  • Hands-on scripting and automation experience usingย Python and Bash.
  • Strong experience withย Prometheus and Grafanaย for monitoring and observability.
  • Experience withย OpenTelemetry, Thanos, logging platforms, or similar observability technologies.
  • Strong understanding ofย SLIs, SLOs, error budgets, incident management, and production troubleshooting.
  • Good understanding of DNS, TCP/IP networking, TLS, VPNs, load balancing, firewalls, and cloud connectivity.
  • Experience working with hybrid or multi-cloud infrastructure, preferably includingย AWS and IBM Cloud.
  • Strong understanding of containers, distributed systems, scalability, availability, and fault tolerance.
  • Experience with disaster recovery, capacity planning, and production resilience.
  • Exposure to regulated or compliance-driven environments such asย HIPAA, SOC 2, PCI DSS, or ISO 27001.
  • Strong analytical, troubleshooting, and root-cause analysis skills.
  • Excellent communication and collaboration skills.
  • Ability to take ownership of critical production systems and operate effectively during high-severity incidents.
  • Experience mentoring engineers and contributing to technical architecture and reliability standards.

Similar jobs