AI-Enabled Platform/SRE Engineer - HYBRID
Why This Role Stands Out
Advance your career by leveraging cutting-edge AI to enhance platform reliability and automate operations in a dynamic hybrid role. You'll thrive here if you're a skilled SRE eager to build innovative solutions and contribute to a forward-thinking technology group. Apply to join this exciting team and gain invaluable experience in a high-growth environment.
Quick Overview
Job Description
Location: Dallas, TX / Scottsdale, AZ (Hybrid)
**Need only local candidates.**
Length of contract : 1 year (and will be renewed after that)
**ROPES ASSESSMENT IS REQUIRED**
**INTERVIEW WILL BE ONSITE IN DALLAS, TX or SCOTTSDALE, AZ**
About the Role
We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.
Key Responsibilities
· Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
· Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
· Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
· Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
· Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
· Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
· Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Core Technical Skills
· Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
· Kubernetes Platform Engineering – 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
· Cloud & Infrastructure Automation – Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
· Software Development – 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
· Observability & Monitoring – Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.
· API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
· AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
Skills
Similar jobs
SRE - Observability
Selby Jennings · Austin, United States
30 minutes agoDevOps engineer 9+yrs(W2 Only )
Cloudberyl LLC · Austin, United States
2 hours agoSite Reliability Engineer (SRE)
Info Way Solutions · United States
2 hours agoDevOps & Platform Engineer
Negocios IT Solutions (P) LTD · Jersey City, United States
2 hours agoMid-Level AWS DevOps Engineer [$301k/yr+] TS/SCI with Security Clearance
SYSTOLIC · Annapolis Junction, United States
2 hours ago$301k/yrAWS/EKS Platform Engineer-Hybrid/Reston VA.-Final interview is F2F in Reston VA.
Elite Technical · Reston, United States
2 hours ago$100/hr