AI-Enabled Platform/SRE Engineer (Local TX)
Quick Overview
Job Description
Role: AI-Enabled Platform/SRE Engineer
Location: Hybrid in Richardson - need local from TX !
Duration: 12 months + extensions
Visa - OPT/ead - C2C
About the Role
We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.
Key Responsibilities
• Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
• Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
• Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
• Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
• Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
• Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
• Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Core Technical Skills
• Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
• Kubernetes Platform Engineering – 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
• Cloud & Infrastructure Automation – Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
• Software Development – 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
• Observability & Monitoring – Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.
• API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
• AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
Skills
Similar jobs
Software Engineer 1 (DevOps) with Security Clearance
Avid Technology Professionals · Linthicum Heights, United States
20 minutes agoDevOps SME
Leidos · Falls Church, United States
36 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Sterling, United States
36 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Springfield, United States
36 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Reston, United States
36 minutes ago$154.1k - $278.5k/yrDevOps SME
Leidos · Chantilly, United States
36 minutes ago$154.1k - $278.5k/yr