Why This Role Stands Out
Elevate your career in a hybrid role where you'll design and operate enterprise observability platforms, leveraging your expertise in Splunk, Elasticsearch, and distributed tracing to build robust solutions. This position is ideal for experienced SREs who thrive on complex technical challenges and are eager to contribute to a reputable company. Apply now to join a forward-thinking team and make a significant impact.
Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
3 weeks ago
DockerRubyAWSELKService MeshSplunkAnsibleAzureBashConsulGoogle CloudGrafanaKafkaKibanaKubernetesPrometheusPythonTerraform
Job Description
Role :: Senior/Lead Site Reliability Engineer Observability
Location :: Remote
Type :: Fulltime
Job Description
Must Have Technical/Functional Skills:
• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
• Hands-on experience administering Splunk Enterprise or Splunk Cloud.
• Strong knowledge of Splunk SPL.
• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
• Experience implementing metrics, logs, and traces as part of a modern observability strategy.
• Experience with Terraform and Infrastructure as Code.
• Programming experience in Python, Go, Ruby, or Bash.
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
• Experience supporting FedRAMP or regulated environments.
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Roles & Responsibilities:
• Design, deploy, and operate enterprise observability platforms.
• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
• Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
• Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
• Automate infrastructure using Terraform and configuration management tools.
Nice to have skills:
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
• Experience supporting FedRAMP or regulated environments.
In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Security Clearance which requires that the person ====, a U.S. Permanent Resident (i.e., a “”), or a Political Asylee or Refugee.
Must Have Technical/Functional Skills:
• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
• Hands-on experience administering Splunk Enterprise or Splunk Cloud.
• Strong knowledge of Splunk SPL.
• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
• Experience implementing metrics, logs, and traces as part of a modern observability strategy.
• Experience with Terraform and Infrastructure as Code.
• Programming experience in Python, Go, Ruby, or Bash.
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
• Experience supporting FedRAMP or regulated environments.
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Roles & Responsibilities:
• Design, deploy, and operate enterprise observability platforms.
• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
• Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
• Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
• Automate infrastructure using Terraform and configuration management tools.
Nice to have skills:
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
• Experience supporting FedRAMP or regulated environments.
In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Security Clearance which requires that the person ====, a U.S. Permanent Resident (i.e., a “”), or a Political Asylee or Refugee.
Similar jobs
- AF
Cloud DevOps Engineer with Security Clearance
NewAccenture Federal Services
Colorado Springs, CO🇺🇸$100.2k - $203.4k/yrOn-siteYesterdayDockerAWSEncryption+16Technology - MR
Senior DevOps Engineer / Lehi / Hybrid
NewMotion Recruitment Partners, LLC
Salt Lake City, UT🇺🇸$130k - $150k/yrHybrid8 hours agoAWSBashGrafana+5Technology - CA
Senior Cloud DevOps Engineer with Security Clearance
CACI
San Antonio, TX🇺🇸$85.8k - $180.2k/yrHybrid2 weeks agoGCPSQLAWS+13Technology - IN
Site Reliability Engineer II
NewInnova
Charlotte, NC🇺🇸$57/hrHybrid8 hours agoAWSTechnology - ES
Cloud Database Platform Engineer with Security Clearance
NewEchelon Services, LLC
Clayton, OH🇺🇸RemoteYesterdayOracleSQLSQL Server+8Technology - PS
Sr. IT DevOps Engineer
NewPrecision System Design Inc.
Cleveland, OH🇺🇸Remote2 days agoAWSAzureGit+3Technology