← Back to Jobs
Technology
GA
DevOps / Site Reliability Engineer (SRE)
Get A WhizAtlanta, GA🇺🇸United StatesPosted 19 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
DevOps / Site Reliability Engineer (SRE) - Multiple Positions
Introduction
We are looking for experienced DevOps / Site Reliability Engineers (SREs) to help build and support cloud infrastructure used by modern AI applications. The person in this role will work with AWS, Kubernetes, CI/CD pipelines, monitoring tools, and automation. You will also work closely with Engineering and Security teams to fix infrastructure issues, improve system reliability, and address security vulnerabilities.
Responsibilities:
- Build and maintain cloud infrastructure using AWS.
- Support production environments running on Kubernetes / Amazon EKS.
- Build and maintain Jenkins CI/CD pipelines.
- Automate infrastructure setup, deployments, and routine operational tasks.
- Provide Tier 2 production support and troubleshoot infrastructure issues.
- Help improve the reliability, availability, and performance of cloud environments.
- Monitor systems and applications using Splunk, Dynatrace, and Grafana.
- Write Python scripts to automate repetitive tasks and improve operations.
- Identify and fix security vulnerabilities in cloud and infrastructure environments.
- Work with Security and Engineering teams to address vulnerabilities found by Mythos AI or similar security tools.
- Help improve monitoring, alerting, logging, and overall system visibility.
- Troubleshoot production incidents and help identify the root cause.
- Support infrastructure used by AI/ML applications.
- Look for opportunities to automate manual processes and improve the overall platform.
Required Skills:
- 6+ years of devops experience.
- Hands-on experience with AWS.
- Strong production experience with Kubernetes / EKS.
- Experience with Jenkins and CI/CD pipelines.
- Experience building and automating cloud infrastructure.
- Experience providing Tier 2 infrastructure or production support.
- Good understanding of system reliability and high availability.
- Experience with Splunk, Dynatrace, or Grafana.
- Strong Python scripting skills.
- Experience fixing security vulnerabilities and hardening infrastructure.
- Good troubleshooting and problem-solving skills.
- Ability to work with Engineering, Security, and Operations teams.
Nice to Have:
- Experience supporting AI/ML platforms or workloads.
- Knowledge of AI infrastructure or Generative AI environments.
- Experience with Mythos AI or similar AI-based security tools.
- Experience with Terraform, Ansible, Docker, Helm, Git, or Argo CD.
- Understanding of SRE concepts such as SLI, SLO, and error budgets.
Skills
Docker
AWS
Splunk
Ansible
Generative AI
Git
Grafana
Helm
Jenkins
Kubernetes
Python
Terraform
Similar jobs
NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems)
One IT Corp · United States
3 minutes agoCloud DevOps Engineer, Remote - 69750
PRIMUS Global Services Inc. · United States
3 minutes agoSite Reliability Engineer - Remote
DivIHN Integration Inc. · United States
3 minutes agoDeVops/DevSecOps - AWS
Delviom LLC · San Antonio, United States
4 minutes agoDevOps Test Engineer with Security Clearance
T-Rex Solutions LLC · Fort Meade, United States
5 minutes ago$160k - $200k/yrDevOps Engineer - Dallas, TX, Austin, TX, Houston, TX, San Antonio, TX.
TechniPros, LLC · Dallas, United States
5 minutes ago