Haystack
← Back to Jobs
Technology

AI-Enabled Platform/SRE Engineer - HYBRID

Excellent Pro Group Inc.Dallas, TX🇺🇸United StatesPosted 28 Jul 2026

Why This Role Stands Out

Advance your career by leveraging cutting-edge AI to enhance platform reliability and automate operations in a dynamic hybrid role. You'll thrive here if you're a skilled SRE eager to build innovative solutions and contribute to a forward-thinking technology group. Apply to join this exciting team and gain invaluable experience in a high-growth environment.

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Location:  Dallas, TX / Scottsdale, AZ (Hybrid)

**Need only local candidates.**

Length of contract : 1 year (and will be renewed after that)

 

**ROPES ASSESSMENT IS REQUIRED**

**INTERVIEW WILL BE ONSITE IN DALLAS, TX or SCOTTSDALE, AZ**

 

About the Role

 

We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.

 

Key Responsibilities

·  Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.

·  Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.

·  Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.

·  Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.

·  Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.

·  Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.

·  Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

Core Technical Skills

·  Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.

·  Kubernetes Platform Engineering – 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.

·  Cloud & Infrastructure Automation – Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.

·  Software Development – 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).

·  Observability & Monitoring – Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.

·  API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.

·  AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.

Skills

Microservices
Node.js
Splunk
Datadog
Generative AI
Google Cloud
Grafana
GraphQL
Helm
Java
Kubernetes
Python
REST
Terraform

Similar jobs