← Back to Jobs
Technology
Site Reliability Engineer (SRE)
Info Way SolutionsUnited States🇺🇸United StatesPosted 28 Jul 2026
Why This Role Stands Out
This hybrid Lead Site Reliability Engineer role offers significant impact by driving enterprise-scale observability platforms using cutting-edge tools like Splunk, ELK, and Kubernetes. You'll thrive here if you possess deep expertise in SRE, DevOps, and cloud technologies, with ample opportunities for skill development in a reputable company. Apply to shape the future of observability and enjoy a flexible work environment.
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
We are looking for an experienced Lead Site Reliability Engineer (SRE) – Observability to join our team and drive the design, implementation, and support of enterprise-scale observability platforms. The ideal candidate will have strong expertise in Splunk, Elasticsearch (ELK), Grafana, Prometheus, OpenTelemetry, Kafka, Terraform, and Kubernetes, with a solid background in Site Reliability Engineering, DevOps, and cloud technologies.
Required Skills
- 7+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, or DevOps.
- Hands-on experience with Splunk Enterprise and/or Splunk Cloud administration.
- Strong proficiency in Splunk SPL.
- Experience with Elasticsearch (ELK Stack), Kibana, Prometheus, Grafana, Grafana Tempo, and OpenTelemetry.
- Experience implementing distributed tracing, monitoring, logging, and alerting solutions.
- Strong knowledge of Kafka and observability pipelines.
- Hands-on experience with Terraform and Infrastructure as Code (IaC).
- Experience with Kubernetes, Docker, and Linux environments.
- Strong scripting skills using Python, Go, Ruby, or Bash.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
Responsibilities
- Design, deploy, and maintain enterprise observability platforms.
- Administer Splunk infrastructure, including Search Head Clusters, Indexers, Heavy Forwarders, and Deployment Servers.
- Build and manage Elasticsearch clusters for large-scale log analytics.
- Develop dashboards, alerts, and monitoring solutions using Splunk, Grafana, and Kibana.
- Implement distributed tracing using OpenTelemetry and Grafana Tempo.
- Automate infrastructure deployments using Terraform.
- Troubleshoot production issues and improve platform reliability, scalability, and performance.
- Collaborate with development and infrastructure teams to enhance monitoring and operational excellence.
Preferred Qualifications
- Splunk Certification.
- Experience with Ansible, Consul, CI/CD pipelines, and Service Mesh technologies.
- Experience working in FedRAMP or other regulated environments.
Skills
Docker
Ruby
AWS
ELK
Service Mesh
Splunk
Ansible
Azure
Bash
Consul
Google Cloud
Grafana
Kafka
Kibana
Kubernetes
Prometheus
Python
Terraform
Similar jobs
SRE - Observability
Selby Jennings · Austin, United States
40 minutes agoDevOps engineer 9+yrs(W2 Only )
Cloudberyl LLC · Austin, United States
2 hours agoAI-Enabled Platform/SRE Engineer - HYBRID
Excellent Pro Group Inc. · Dallas, United States
2 hours agoDevOps & Platform Engineer
Negocios IT Solutions (P) LTD · Jersey City, United States
2 hours agoMid-Level AWS DevOps Engineer [$301k/yr+] TS/SCI with Security Clearance
SYSTOLIC · Annapolis Junction, United States
2 hours ago$301k/yrAWS/EKS Platform Engineer-Hybrid/Reston VA.-Final interview is F2F in Reston VA.
Elite Technical · Reston, United States
2 hours ago$100/hr