Haystack
← Back to Jobs
Technology
AI

Site Reliability Engineer

ARK Infotech SpectrumAtlanta, GA🇺🇸United StatesPosted 1 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Atlanta, GA, United States
Posted
Yesterday
DockerMicroservicesShellSpringSpring BootAWSELKSplunkAnsibleAzureDatadogGitHub ActionsGitLab CIGoogle CloudGrafanaHelmJavaJenkinsKafkaKubernetesPrometheusRESTTerraform

Job Description

Job Title: Site Reliability Engineer (SRE)
Location: Atlanta, Georgia (Hybrid)
Role Summary

We are seeking an experienced Technology Consultant Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability, Java, and production reliability. The ideal candidate will have experience supporting highly available and distributed enterprise applications, troubleshooting complex production issues, and driving automation and reliability improvements.

The role requires close collaboration with application engineering, DevOps, cloud, infrastructure, and support teams to improve application availability, scalability, performance, and operational efficiency.

Day to Day Job Duties

  • Manage and support business-critical applications running on Kubernetesand containerized platforms.
  • Monitor application and platform health and proactively identify reliability, availability, and performance issues.
  • Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
  • Implement and enhance observability solutionscovering metrics, logs, traces, dashboards, and alerting.
  • Support and troubleshoot Java/Spring Boot and Microservices-based applications.
  • Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
  • Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
  • Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
  • Participate in incident, problem, change, and production release management activities.
  • Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
  • Support CI/CD pipelines and improve application deployment and release processes.
  • Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
  • Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.

Basic Qualifications

  • 6+ yearsof experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ yearsof hands-on experience with Kubernetes, Docker, and containerized application environments.
  • 4+ yearsof experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ yearsof experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.

Nice to Have

  • Strong understanding of SLI, SLO, SLA, Error Budgeting, and SRE principles.
  • Experience with Kubernetes deployment and troubleshooting tools such as Helm.
  • Experience with AWS, Azure, or Google Cloud Platform.
  • Knowledge of Linux/Unix and Shell scripting.
  • Experience with Kafka, IBM MQ, or other messaging technologies.
  • Knowledge of Terraform, Ansible, or other Infrastructure as Code tools.
  • Experience with Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
  • Experience implementing distributed tracing and application performance monitoring.
  • Knowledge of incident management and ITIL processes.
  • Experience supporting high-volume, highly available, distributed enterprise applications.
  • Strong analytical, troubleshooting, communication, and problem-solving skills

Similar jobs