Quick Overview
Job Description
K&K Global Talent Solutions Inc. is an international recruiting agency that has been providing technical resources in the Canada and the USA region since 1993.
This position is with one of our clients in USA, who is actively hiring candidates to expand their teams.
Role:- Senior/Lead Site Reliability Engineer - Observability
Location:- Remote (USA)
Fulltime
Job Description
Must Have Technical/Functional Skills:
• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
• Hands-on experience administering Splunk Enterprise or Splunk Cloud.
• Strong knowledge of Splunk SPL.
• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
• Experience implementing metrics, logs, and traces as part of a modern observability strategy.
• Experience with Terraform and Infrastructure as Code.
• Programming experience in Python, Go, Ruby, or Bash.
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
• Experience supporting FedRAMP or regulated environments.
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Roles & Responsibilities:
• Design, deploy, and operate enterprise observability platforms.
• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
• Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
• Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
• Automate infrastructure using Terraform and configuration management tools.
Nice to have skills:
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
• Experience supporting FedRAMP or regulated environments.
Similar jobs
- AG
Google Cloud Platform AI Platform Engineer
NewASCII Group LLC
Scottsdale, AZ🇺🇸$65/hrOn-site19 hours agoGenerative AIGoogle CloudGrafana+4Technology - IT
Platform DevOps Engineer - Linux 102423 with Security Clearance
NewInformation Technology Engineering Corporation
Aurora, CO🇺🇸Hybrid19 hours agoDockerEncryptionAgile+8Technology - PC
DevOps/Platform Engineer (Kubernetes/ API & AI Gateway)
Pyramid Consulting, Inc.
Minneapolis, MN🇺🇸$65 - $68/hrOn-site2 weeks agoAPI GatewayAWSEnvoy+11Technology - SA
AWS DevSecOps Platform Engineer
NewSapear Inc
St. Louis, MO🇺🇸On-site19 hours agoDockerAWSLoad Balancing+18Technology - IT
Data Platform Engineer
NewISite Technologies Inc
Virginia Beach, VA🇺🇸Hybrid19 hours agoSQLAzureTechnology - QU
Platform Engineer Infrastructure DevOps
NewQTech US Inc
Denver, CO🇺🇸On-site19 hours agoTechnology