Quick Overview
Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
Yesterday
DockerRubyAWSELKService MeshSplunkAnsibleAzureBashConsulGoogle CloudGrafanaKafkaKibanaKubernetesPrometheusPythonTerraform
Job Description
Job Role: Senior/Lead Site Reliability Engineer - Observability
Location: Remote
Job Description:
Must Have Technical/Functional Skills
- 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
- Hands-on experience administering Splunk Enterprise or Splunk Cloud.
- Strong knowledge of Splunk SPL.
- Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, Open Telemetry, and Kafka.
- Experience implementing metrics, logs, and traces as part of a modern observability strategy.
- Experience with Terraform and Infrastructure as Code.
- Programming experience in Python, Go, Ruby, or Bash.
- Splunk certification.
- Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
- Experience supporting FedRAMP or regulated environments.
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, Open Telemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Roles & Responsibilities:
- Design, deploy, and operate enterprise observability platforms.
- Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
- Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
- Design, deploy, and support distributed tracing platforms using Grafana Tempo and Open Telemetry.
- Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
- Scale Prometheus, Grafana, Kafka, Tempo, and Open Telemetry-based monitoring solutions.
- Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
- Automate infrastructure using Terraform and configuration management tools.
Nice to have skills:
- Splunk certification.
- Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
- Experience supporting FedRAMP or regulated environments.
Similar jobs
- WS
DevOps Cloud Engineer with Security Clearance
White Sky Technologies
Annapolis Junction, MD🇺🇸$140k - $235k/yrHybrid5 weeks agoDockerRubyAWS+8Technology - WS
Senior DevOps Cloud Engineer with Security Clearance
White Sky Technologies
Annapolis Junction, MD🇺🇸$140k - $235k/yrHybrid6 weeks agoDockerRubyAWS+8Technology - EI
SRE PRODUCTION SUPPORT
NewEcho IT Solutions, Inc.
Alpharetta, GA🇺🇸On-siteYesterdayManufacturing - AS
AWS Lead DevOps Engineer
Apex Systems
Plano, TX🇺🇸On-site1 week agoMicroservicesAWSSnowflake+1Technology - RA
Web Application / DevOps Architect Onsite in San Jose, CA
NewRapidIT, Inc
San Jose, CA🇺🇸HybridYesterdayDockerFastAPIMemcached+16Technology - IC
Site Reliability Engineer (SRE)
Infinite Computer Solutions (ICS)
Frisco, TX🇺🇸Hybrid5 weeks agoDockerShellSplunk+11Technology