Haystack
← Back to Jobs
Technology
TO

SRE Platform Engineer AI Agents

TOPSYSITAtlanta, GA🇺🇸United StatesPosted Sep 30, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Atlanta, GA, United States
Posted
Yesterday
JavaKubernetesPython

Job Description

SRE Platform Engineer AI Agents
Job Title: SRE Platform Engineer AI Agents
Experience: 8+ Years
Location: Atlanta, GA (Inperson Interview)
Job Summary
We are looking for an experienced SRE Platform Engineer with 8+ years of experience in Site Reliability Engineering, platform engineering, cloud infrastructure, Kubernetes, and automation. The ideal candidate will also have hands-on exposure to AI Agents / Agentic AI and be comfortable working on modern AI-driven platform solutions.
The candidate should have strong knowledge of SRE principles, SLA/SLO, Kubernetes, cloud platforms, automation, monitoring, and production support, along with solid programming and problem-solving skills.
Key Responsibilities
  • Design, build, maintain, and improve highly reliable and scalable platform infrastructure.
  • Implement Site Reliability Engineering (SRE) practices across production environments.
  • Define and monitor SLAs, SLOs, SLIs, error budgets, and service reliability metrics.
  • Develop and maintain Kubernetes-based applications and platform infrastructure.
  • Troubleshoot production issues, perform root-cause analysis, and implement long-term corrective actions.
  • Build automation for deployment, monitoring, infrastructure management, and operational processes.
  • Work with AI Agents / Agentic AI solutions and integrate AI capabilities into platform and operational workflows.
  • Support CI/CD pipelines and DevOps automation.
  • Monitor application and infrastructure health using logging, metrics, and alerting tools.
  • Collaborate with development, DevOps, cloud, and AI engineering teams.
  • Participate in technical design discussions and contribute to platform architecture.
  • Write clean, efficient code/scripts and demonstrate strong problem-solving abilities.
Required Skills
  • 8+ years of experience in SRE, Platform Engineering, DevOps, or related infrastructure roles.
  • Strong understanding of SRE concepts and practices.
  • Strong knowledge of SLA, SLO, SLI, and Error Budgets.
  • Hands-on experience with Kubernetes and containerized environments.
  • Strong understanding of cloud infrastructure and production systems.
  • Experience with CI/CD, automation, monitoring, logging, and incident management.
  • Strong programming/scripting skills in languages such as Python, Java, Go, or similar.
  • Experience with AI Agents / Agentic AI or AI-driven automation.
  • Strong troubleshooting, debugging, and analytical skills.
  • Excellent communication and collaboration skills.
Interview Focus Areas
Candidates should be prepared to discuss:
  • Current/recent project architecture, responsibilities, and technical contributions.
  • SRE concepts: SLA, SLO, SLI, error budgets, reliability, and incident management.
  • Kubernetes fundamentals: pods, deployments, services, namespaces, scaling, configuration, and troubleshooting.
  • Programming/coding problems, including exercises such as palindrome-number logic.
  • Practical problem-solving and debugging scenarios.
  • Experience designing or implementing AI Agents / Agentic AI solutions.

Similar jobs