Why This Role Stands Out
Leverage your extensive SRE experience to drive cutting-edge observability solutions in a fully remote environment at Vaarida Technologies. This role offers significant growth potential as you implement advanced metrics, logs, and tracing strategies, making it ideal for seasoned professionals passionate about building robust, scalable systems. Apply today to shape the future of reliability!
Quick Overview
Job Description
Position - Senior/Lead Site Reliability Engineer Observability
Location - 100% Remote
Experience - 8+ Years
Type - Full Time
Technology Stack - Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Job Description -
Must Have Technical/Functional Skills:
7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
Hands-on experience administering Splunk Enterprise or Splunk Cloud.
Strong knowledge of Splunk SPL.
Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
Experience implementing metrics, logs, and traces as part of a modern observability strategy.
Experience with Terraform and Infrastructure as Code.
Programming experience in Python, Go, Ruby, or Bash.
Splunk certification.
Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
Experience supporting FedRAMP or regulated environments.
Roles & Responsibilities:
Design, deploy, and operate enterprise observability platforms.
Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
Automate infrastructure using Terraform and configuration management tools.
Nice to have skills:
Splunk certification.
Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
Experience supporting FedRAMP or regulated environments.
Similar jobs
- RI
Devsecops Engineer
NewRIIASH LLC
United States🇺🇸HybridYesterdayEngineering - EP
AppDynamics Platform Engineer
Newe-IT Professionals Corp.
United States🇺🇸HybridYesterdayAWSSplunkDatadogTechnology - NI
DevOps Engineer
NewNightwing
Arlington, Virginia🇺🇸Hybrid25 minutes agoDockerAWSELK+13Technology - KA
Lead Platform Engineer
NewKainos
United States🇺🇸Hybrid25 minutes agoAWSAgileAzure+1Technology - BL
Director, Principal Platform Engineer (Compute)
NewBlackRock
New York City, New York🇺🇸$215k - $275k/yrHybrid27 minutes agoAWSAgileAzure+4Technology - EP
Site Reliability Engineer - Greenwood Village, Colorado (Hybrid) - W2 role
NewEmpower Professionals
Greenwood Village, CO🇺🇸HybridYesterdayNode.jsSQLAWS+14Technology