Haystack
← Back to Jobs
Technology
II

SRE Engineer

Icon International Group LLCDallas, TX🇺🇸United StatesPosted 8 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Dallas, TX, United States
Posted
19 hours ago
MicroservicesNode.jsSplunkDatadogGenerative AIGoogle CloudGrafanaGraphQLHelmJavaKubernetesPythonRESTTerraform

Job Description

Role: SRE Engineer
Duration: 12 Months 
Work Mode: Hybrid

Location: Scottsdale, Arizona 85260/ Dallas, TX

Position Overview:
We are seeking a Senior Kubernetes-focused Site Reliability Engineer (SRE) with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale.

Key Responsibilities
  • Software Development & Tooling: Build automation and operational tools using Python, Java, and Node.js to improve efficiency, scalability, and platform operations.
  • AI-Driven Operations (AIOps): Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
  • API & Microservices Reliability: Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
  • Kubernetes Platform Management: Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • High Availability & Disaster Recovery: Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
  • Observability & Monitoring: Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
  • Operational Excellence: Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Core Technical Skills
  • Site Reliability Engineering (SRE): Deep understanding of reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
  • Kubernetes Platform Engineering: 5+ years of strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
  • Software Development: 5+ years of advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
  • Cloud & Infrastructure Automation: Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
  • AI-Driven Operations (AIOps): Hands-on experience applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, operational workflow automation, and runbook execution.
  • API & Microservices Engineering: Expertise in Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
  • Observability & Monitoring: Proficiency in Splunk, Grafana, Datadog, AppDynamics, alerting design, and platform health monitoring.
 

Similar jobs