Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Atlanta, GA, United States
Posted
Yesterday
JavaKubernetesPython
Job Description
SRE Platform Engineer AI Agents
Job Title: SRE Platform Engineer AI Agents
Experience: 8+ Years
Location: Atlanta, GA (Inperson Interview)
Experience: 8+ Years
Location: Atlanta, GA (Inperson Interview)
Job Summary
We are looking for an experienced SRE Platform Engineer with 8+ years of experience in Site Reliability Engineering, platform engineering, cloud infrastructure, Kubernetes, and automation. The ideal candidate will also have hands-on exposure to AI Agents / Agentic AI and be comfortable working on modern AI-driven platform solutions.
The candidate should have strong knowledge of SRE principles, SLA/SLO, Kubernetes, cloud platforms, automation, monitoring, and production support, along with solid programming and problem-solving skills.
Key Responsibilities
- Design, build, maintain, and improve highly reliable and scalable platform infrastructure.
- Implement Site Reliability Engineering (SRE) practices across production environments.
- Define and monitor SLAs, SLOs, SLIs, error budgets, and service reliability metrics.
- Develop and maintain Kubernetes-based applications and platform infrastructure.
- Troubleshoot production issues, perform root-cause analysis, and implement long-term corrective actions.
- Build automation for deployment, monitoring, infrastructure management, and operational processes.
- Work with AI Agents / Agentic AI solutions and integrate AI capabilities into platform and operational workflows.
- Support CI/CD pipelines and DevOps automation.
- Monitor application and infrastructure health using logging, metrics, and alerting tools.
- Collaborate with development, DevOps, cloud, and AI engineering teams.
- Participate in technical design discussions and contribute to platform architecture.
- Write clean, efficient code/scripts and demonstrate strong problem-solving abilities.
Required Skills
- 8+ years of experience in SRE, Platform Engineering, DevOps, or related infrastructure roles.
- Strong understanding of SRE concepts and practices.
- Strong knowledge of SLA, SLO, SLI, and Error Budgets.
- Hands-on experience with Kubernetes and containerized environments.
- Strong understanding of cloud infrastructure and production systems.
- Experience with CI/CD, automation, monitoring, logging, and incident management.
- Strong programming/scripting skills in languages such as Python, Java, Go, or similar.
- Experience with AI Agents / Agentic AI or AI-driven automation.
- Strong troubleshooting, debugging, and analytical skills.
- Excellent communication and collaboration skills.
Interview Focus Areas
Candidates should be prepared to discuss:
- Current/recent project architecture, responsibilities, and technical contributions.
- SRE concepts: SLA, SLO, SLI, error budgets, reliability, and incident management.
- Kubernetes fundamentals: pods, deployments, services, namespaces, scaling, configuration, and troubleshooting.
- Programming/coding problems, including exercises such as palindrome-number logic.
- Practical problem-solving and debugging scenarios.
- Experience designing or implementing AI Agents / Agentic AI solutions.
Similar jobs
- RA
Site Reliability Engineer
NewRaydar
United States🇺🇸Remote23 hours agoNode.jsHelmKubernetes+3Technology - RL
Lead ForgeRock Platform Engineer (IAM)
NewResource Logistics
San Antonio, TX🇺🇸On-siteYesterdayMicroservicesSpringSpring Boot+14Technology - GE
Site Reliability Engineer
NewGenesis10
Jersey City, NJ🇺🇸$60 - $68/hrHybridYesterdayLoad BalancingSplunkAzure+10Technology - ES
Senior Data DevOps Engineer with Azure
EPAM Systems
United States🇺🇸Hybrid5 days agoMLOpsAzureData Pipeline+1Technology - CD
Network Automation Engineer
NewCloud Destinations LLC
United States🇺🇸RemoteYesterdayDjangoDockerTDD+11Technology - QS
Palantir Platform Engineer with Security Clearance
NewQuantum Science Solutions
Herndon, VA🇺🇸HybridYesterdayETLPythonTechnology