Haystack
← Back to Jobs
Technology
SS

Lead Site Reliability Engineer

SIGNin Solutions Inc.United States🇺🇸United StatesPosted Oct 7, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
DockerRubyAWSELKSplunkAnsibleAzureBashConsulGoogle CloudGrafanaKafkaKibanaKubernetesPrometheusPythonTerraform

Job Description

Key Responsibilities
  • Design, deploy, and operate enterprise observability platforms.
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search
Head Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and
OpenTelemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention
strategies.
  • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana,
Kibana, and Tempo.
  • Automate infrastructure using Terraform and configuration management tools.

Required Qualifications
  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing,
OpenTelemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.

Preferred Qualifications
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service
mesh technologies.
  • Experience supporting FedRAMP or regulated environments.
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana
Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python,
Go, Ruby, Bash, AWS, Ansible, Consul.
The successful applicant may perform work in FedRAMP High or IL-5 environments
Keywords
SRE, Site Reliability Engineering, Observability, Splunk, Splunk Enterprise, Splunk Cloud,
Elasticsearch, ELK, Kibana, Prometheus, Grafana, Tempo, Grafana Tempo, Distributed Tracing,
OpenTelemetry, Traces, Kafka, Terraform, Kubernetes, Linux, DevOps.

Similar jobs