Haystack
← Back to Jobs
Technology
GA

DevOps / Site Reliability Engineer (SRE)

Get A WhizAtlanta, GA🇺🇸United StatesPosted 19 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

DevOps / Site Reliability Engineer (SRE) - Multiple Positions

Introduction

We are looking for experienced DevOps / Site Reliability Engineers (SREs) to help build and support cloud infrastructure used by modern AI applications. The person in this role will work with AWS, Kubernetes, CI/CD pipelines, monitoring tools, and automation. You will also work closely with Engineering and Security teams to fix infrastructure issues, improve system reliability, and address security vulnerabilities.

Responsibilities:

  • Build and maintain cloud infrastructure using AWS.
  • Support production environments running on Kubernetes / Amazon EKS.
  • Build and maintain Jenkins CI/CD pipelines.
  • Automate infrastructure setup, deployments, and routine operational tasks.
  • Provide Tier 2 production support and troubleshoot infrastructure issues.
  • Help improve the reliability, availability, and performance of cloud environments.
  • Monitor systems and applications using Splunk, Dynatrace, and Grafana.
  • Write Python scripts to automate repetitive tasks and improve operations.
  • Identify and fix security vulnerabilities in cloud and infrastructure environments.
  • Work with Security and Engineering teams to address vulnerabilities found by Mythos AI or similar security tools.
  • Help improve monitoring, alerting, logging, and overall system visibility.
  • Troubleshoot production incidents and help identify the root cause.
  • Support infrastructure used by AI/ML applications.
  • Look for opportunities to automate manual processes and improve the overall platform.

Required Skills:

  • 6+ years of devops experience.
  • Hands-on experience with AWS.
  • Strong production experience with Kubernetes / EKS.
  • Experience with Jenkins and CI/CD pipelines.
  • Experience building and automating cloud infrastructure.
  • Experience providing Tier 2 infrastructure or production support.
  • Good understanding of system reliability and high availability.
  • Experience with Splunk, Dynatrace, or Grafana.
  • Strong Python scripting skills.
  • Experience fixing security vulnerabilities and hardening infrastructure.
  • Good troubleshooting and problem-solving skills.
  • Ability to work with Engineering, Security, and Operations teams.

Nice to Have:

  • Experience supporting AI/ML platforms or workloads.
  • Knowledge of AI infrastructure or Generative AI environments.
  • Experience with Mythos AI or similar AI-based security tools.
  • Experience with Terraform, Ansible, Docker, Helm, Git, or Argo CD.
  • Understanding of SRE concepts such as SLI, SLO, and error budgets.

Skills

Docker
AWS
Splunk
Ansible
Generative AI
Git
Grafana
Helm
Jenkins
Kubernetes
Python
Terraform

Similar jobs