Haystack
← Back to Jobs
Remote
Technology

Lead Site Reliability Engineer – Observability (US Remote)||W2

Ukund IncUnited States🇺🇸United StatesPosted 24 Jul 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Lead Site Reliability Engineer – Observability (US Remote)

Location: Remote, prefer PST hours

Key Responsibilities

  • Design, deploy, and operate enterprise observability platforms.

  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.

  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.

  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.

  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.

  • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.

  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.

  • Automate infrastructure using Terraform and configuration management tools.

 

Required Qualifications

  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.

  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.

  • Strong knowledge of Splunk SPL.

  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.

  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.

  • Experience with Terraform and Infrastructure as Code.

  • Programming experience in Python, Go, Ruby, or Bash.

 

Preferred Qualifications

  • Splunk certification.

  • Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.

  • Experience supporting FedRAMP or regulated environments.

 

Technology Stack

Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.

 

Skills

Docker
Ruby
AWS
ELK
Service Mesh
Splunk
Ansible
Azure
Bash
Consul
Google Cloud
Grafana
Kafka
Kibana
Kubernetes
Prometheus
Python
Terraform

Similar jobs