Haystack
← Back to Jobs
Technology

AI-Enabled Platform/SRE Engineer- Local TX

TechCafeHub LLCRichardson, TX🇺🇸United StatesPosted 31 Jul 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

AI-Enabled Platform/SRE Engineer- Local TX (f2f) 

Introduction:

As an AI-Enabled Platform/SRE Engineer in our team based in Texas, you will play a crucial role in building and maintaining automation tools, leveraging AI technologies, managing Kubernetes platforms, ensuring high availability, and driving operational excellence. You will work closely with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

Responsibilities:

  • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
  • Leverage AI and Generative AI technologies to automate alert analysis, incident response, operational workflows, and runbook execution.
  • Implement API and microservices reliability solutions using various tools and strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • Ensure platform reliability and high availability through active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
  • Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics.
  • Drive SRE best practices and operational excellence by collaborating with cross-functional teams.

Requirements:

Required Skills:

  • Site Reliability Engineering (SRE) experience
  • 5+ years of hands-on experience with GKE and Rancher RKE2
  • Strong knowledge of Cloud & Infrastructure Automation, including Google Cloud Platform and Terraform
  • Advanced programming skills in Python and Java
  • Experience with observability & monitoring tools
  • Proficiency in API & Microservices Engineering
  • Knowledge of AI-Driven Operations (AIOps)

Preferred Skills:

  • Experience with Continuous Integration/Continuous Delivery (CI/CD)
  • Familiarity with Incident Management and Disaster Recovery practices
  • Understanding of High Availability concepts
  • Knowledge of Kubernetes, Routing, and Operational Excellence

Skills

Microservices
Node.js
Splunk
Datadog
Generative AI
Google Cloud
Grafana
Java
Kubernetes
Python
Terraform

Similar jobs