Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Dallas, TX, United States
Posted
19 hours ago
MicroservicesNode.jsSplunkDatadogGenerative AIGoogle CloudGrafanaGraphQLHelmJavaKubernetesPythonRESTTerraform
Job Description
Role: SRE Engineer
Duration: 12 Months
Work Mode: Hybrid
Location: Scottsdale, Arizona 85260/ Dallas, TX
Position Overview:We are seeking a Senior Kubernetes-focused Site Reliability Engineer (SRE) with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.
Key Responsibilities
- Software Development & Tooling: Build automation and operational tools using Python, Java, and Node.js to improve efficiency, scalability, and platform operations.
- AI-Driven Operations (AIOps): Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
- API & Microservices Reliability: Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
- Kubernetes Platform Management: Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
- High Availability & Disaster Recovery: Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
- Observability & Monitoring: Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
- Operational Excellence: Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
- Site Reliability Engineering (SRE): Deep understanding of reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
- Kubernetes Platform Engineering: 5+ years of strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
- Software Development: 5+ years of advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
- Cloud & Infrastructure Automation: Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
- AI-Driven Operations (AIOps): Hands-on experience applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, operational workflow automation, and runbook execution.
- API & Microservices Engineering: Expertise in Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
- Observability & Monitoring: Proficiency in Splunk, Grafana, Datadog, AppDynamics, alerting design, and platform health monitoring.
Similar jobs
- RA
Staff Platform Engineer
NewRobots and Pencils
US Remote🇺🇸Remote5 hours agoMicroservicesAWSService Mesh+8Technology - MA
Site Reliability Engineer
NewMaintainx
San Francisco🇺🇸Remote5 hours agoNode.jsShellHTTP+2Technology - BL
Sr. Software Engineer, Site Reliability
NewBloomerang
Remote🇺🇸$114.8k - $150k/yrRemote3 hours ago401kNode.jsPHP+9Technology - BF
Senior DevOps Engineer
NewBeyond Finance
Chicago🇺🇸4 hours agoDockerAWSLinear+11Technology - PG
Cloud Platform Engineer Kubernetes, Onsite - 69995
NewPRIMUS Global Services Inc.
Plano, TX🇺🇸On-site19 hours agoShellAWSNginx+9Technology - TA
DevOps Software Engineer, TS/SCI with Poly with Security Clearance
NewTalentOps
Annapolis Junction, MD🇺🇸$160k - $230k/yrRemote19 hours agoDockerMariaDBMySQL+13Technology