Quick Overview
Job Description
Role: Senior DevOps / Site Reliability Engineer (SRE)
Location: Remote
Core Focus
We are seeking a Senior DevOps / Site Reliability Engineer with deep expertise in cloud-native observability, Kubernetes, and modern monitoring platforms. This role will design, implement, and optimize observability solutions across our AWS infrastructure, leveraging the Grafana ecosystem, eBPF, and Kubernetes to improve system reliability, performance, and operational visibility.
5-8 years experience
Key Responsibilities
Design, implement, and maintain scalable observability solutions across cloud-native environments.
Deploy and optimize the Grafana stack, including Grafana Beyla, Grafana Alloy, and OpenTelemetry-based instrumentation.
Leverage eBPF to enable low-overhead application and infrastructure monitoring, tracing, and performance analysis.
Support and optimize Amazon EKS clusters to ensure high availability, scalability, and operational excellence.
Develop monitoring, alerting, and dashboarding solutions that provide actionable insights into platform health.
Partner with engineering teams to improve reliability, incident response, capacity planning, and performance optimization.
Drive SRE and DevOps best practices through automation, operational standards, and continuous improvement.
Required Qualifications
Strong experience as a DevOps Engineer or Site Reliability Engineer supporting production cloud environments.
Deep expertise with Amazon EKS and Kubernetes administration.
Hands-on experience with the Grafana observability stack, including Grafana, Grafana Beyla, and Grafana Alloy.
Strong knowledge of eBPF for application performance monitoring, tracing, and infrastructure observability.
Experience implementing OpenTelemetry and modern observability frameworks.
Proficiency with AWS, Infrastructure as Code, CI/CD, and cloud-native operational practices.
Strong troubleshooting skills with a focus on reliability, scalability, automation, and performance optimization.
Experience in healthcare or other highly regulated environments is a plus.
Similar jobs
- SC
AWS Data- AI Platform Engineer
NewSource Code Technologies LLC
Charlotte, NC🇺🇸Hybrid22 hours agoDynamoDBAPI GatewayAWS+6Technology - LS
Graph AI Platform Engineer
NewLincoln Softtech LLC
Dallas, TX🇺🇸Hybrid22 hours agoNeo4jPyTorchTechnology - MT
SRE & Platform Engineering Lead
NewMicrogreen Technologies LLC
Dallas, TX🇺🇸On-site22 hours agoAWSCDNDNS+3Technology - TG
Infrastructure Operations Engineer
NewTalent Groups
United States🇺🇸Hybrid22 hours agoSplunkAzurePowerShell+2Technology - IA
Senior Site Reliability Engineer
NewIO Associates
Menlo Park, CA🇺🇸Hybrid22 hours agoAWSELKAzure+9Technology - BI
SRE lead @ Phoenix AZ 85054 (Hybrid) - Only on W2
NewBURGEON IT SERVICES LLC
Phoenix, AZ🇺🇸Hybrid22 hours agoScrumGoogle CloudJava+1Technology