Haystack
← Back to Jobs
Technology

RE : Senior Site Reliability Engineer (SRE) – Observability Platform

Info Way SolutionsUnited States🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

About the Role
We are seeking a highly skilled Senior Site Reliability Engineer (SRE) – Observability Platform to design, build, and operate enterprise-scale observability solutions across cloud environments. The ideal candidate will have extensive experience with logging, metrics, distributed tracing, monitoring, and infrastructure automation, helping drive reliability, scalability, and operational excellence for mission-critical applications.
This role requires expertise in modern observability technologies including Splunk, Elasticsearch, Prometheus, Grafana, OpenTelemetry, Grafana Tempo, Kafka, and Terraform.

Key Responsibilities
  • Design, deploy, and manage enterprise observability platforms supporting large-scale cloud infrastructure.
  • Administer and maintain Splunk Enterprise and Splunk Cloud, including:
    • Indexers
    • Search Head Clusters (SHC)
    • Heavy Forwarders
    • Deployment Servers
  • Deploy, optimize, and manage Elasticsearch/ELK clusters for centralized logging and search.
  • Design and support distributed tracing solutions using Grafana Tempo and OpenTelemetry.
  • Build and maintain end-to-end observability pipelines for logs, metrics, and traces.
  • Define instrumentation standards, trace retention policies, and monitoring best practices.
  • Scale and manage monitoring platforms including PrometheusGrafanaKafka, and Tempo.
  • Develop dashboards, alerts, reports, analytics, and trace visualizations using:
    • Splunk SPL
    • Grafana
    • Kibana
    • Tempo
  • Automate infrastructure deployment using Terraform and Infrastructure as Code (IaC).
  • Collaborate with development, DevOps, and platform engineering teams to improve system reliability and performance.
  • Perform troubleshooting, root cause analysis, and incident response for production environments.
  • Ensure high availability, security, scalability, and compliance of observability platforms.

Required Qualifications
  • Bachelor''s degree in Computer Science, Information Technology, or a related field (or equivalent experience).
  • 7+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Cloud Infrastructure.
  • Hands-on administration experience with Splunk Enterprise or Splunk Cloud.
  • Strong expertise in Splunk Search Processing Language (SPL).
  • Experience with:
    • Elasticsearch / ELK Stack
    • Prometheus
    • Grafana
    • Grafana Tempo
    • OpenTelemetry
    • Distributed Tracing
    • Kafka
  • Strong understanding of modern observability practices involving metrics, logs, and traces.
  • Experience with Terraform and Infrastructure as Code (IaC).
  • Proficiency in one or more scripting/programming languages:
    • Python
    • Go
    • Ruby
    • Bash
  • Strong Linux administration and troubleshooting skills.

Preferred Qualifications
  • Splunk Certified Administrator or Splunk Architect certification.
  • Experience with Kubernetes and Docker.
  • Experience with AWS, Azure, or Google Cloud Platform (Google Cloud Platform).
  • Experience using Ansible, Consul, and CI/CD pipelines.
  • Knowledge of Service Mesh technologies (Istio, Linkerd, etc.).
  • Experience supporting FedRAMP High, IL-5, or other regulated environments.
  • Strong understanding of security, compliance, and cloud governance.

Technical Skills
Observability
  • Splunk Enterprise
  • Splunk Cloud
  • Splunk SPL
  • Elasticsearch
  • ELK Stack
  • Kibana
  • Prometheus
  • Grafana
  • Grafana Tempo
  • OpenTelemetry
  • Distributed Tracing
  • Kafka
Cloud & Infrastructure
  • AWS
  • Azure
  • Google Cloud Platform (Google Cloud Platform)
  • Kubernetes
  • Docker
  • Linux
Automation & DevOps
  • Terraform
  • Infrastructure as Code (IaC)
  • Ansible
  • CI/CD Pipelines
  • Git
Programming
  • Python
  • Go
  • Bash
  • Ruby

Preferred Competencies
  • Strong analytical and troubleshooting skills.
  • Experience managing enterprise-scale monitoring platforms.
  • Excellent communication and cross-functional collaboration.
  • Ability to automate operational processes.
  • Strong incident management and root cause analysis capabilities.
  • Experience designing scalable, highly available cloud-native observability solutions.

Eligibility Requirements
  • Must be a U.S. Citizen or U.S. National (U.S. Person).
  • Work must be performed from within the United States.
  • Ability to support environments subject to FedRAMP High and Impact Level 5 (IL-5) security requirements.

Skills

Docker
Ruby
AWS
ELK
Service Mesh
Splunk
Ansible
Azure
Bash
Consul
Git
Google Cloud
Grafana
Istio
Kafka
Kibana
Kubernetes
Prometheus
Python
Terraform

Similar jobs