Quick Overview
Job Description
Job Title: Site Reliability Engineer (SRE)
Location: Atlanta, Georgia (Hybrid)
Role Summary
We are seeking an experienced Technology Consultant Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability, Java, and production reliability. The ideal candidate will have experience supporting highly available and distributed enterprise applications, troubleshooting complex production issues, and driving automation and reliability improvements.
The role requires close collaboration with application engineering, DevOps, cloud, infrastructure, and support teams to improve application availability, scalability, performance, and operational efficiency.
Day to Day Job Duties
- Manage and support business-critical applications running on Kubernetesand containerized platforms.
- Monitor application and platform health and proactively identify reliability, availability, and performance issues.
- Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
- Implement and enhance observability solutionscovering metrics, logs, traces, dashboards, and alerting.
- Support and troubleshoot Java/Spring Boot and Microservices-based applications.
- Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
- Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
- Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
- Participate in incident, problem, change, and production release management activities.
- Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
- Support CI/CD pipelines and improve application deployment and release processes.
- Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
- Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.
Basic Qualifications
- 6+ yearsof experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
- 4+ yearsof hands-on experience with Kubernetes, Docker, and containerized application environments.
- 4+ yearsof experience with Java, Spring Boot, Microservices, and REST APIs.
- 3+ yearsof experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.
Nice to Have
- Strong understanding of SLI, SLO, SLA, Error Budgeting, and SRE principles.
- Experience with Kubernetes deployment and troubleshooting tools such as Helm.
- Experience with AWS, Azure, or Google Cloud Platform.
- Knowledge of Linux/Unix and Shell scripting.
- Experience with Kafka, IBM MQ, or other messaging technologies.
- Knowledge of Terraform, Ansible, or other Infrastructure as Code tools.
- Experience with Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
- Experience implementing distributed tracing and application performance monitoring.
- Knowledge of incident management and ITIL processes.
- Experience supporting high-volume, highly available, distributed enterprise applications.
- Strong analytical, troubleshooting, communication, and problem-solving skills
Similar jobs
- ON
Site Reliability Engineer
NewOnBoard
United States🇺🇸11 hours ago401kSQLAzure+11Technology - SE
Senior Staff Data Platform Engineer - Apache Iceberg - Apache Spark
NewServiceNow
Santa Clara, CALIFORNIA🇺🇸Hybrid14 hours agoTechnology - SE
Senior Staff Data Platform Engineer - Apache Iceberg - Apache Spark
NewServiceNow
San Diego, CALIFORNIA🇺🇸Hybrid14 hours agoTechnology - OO
Site Reliability Engineer
NewOoma
Remote🇺🇸RemoteYesterdayMicroservicesELKLoad Balancing+15Technology - EG
IAM Automation & Security Platform Engineer
NeweTelligent Group LLC
Herndon🇺🇸15 hours agoOAuthAzureCompliance+3Technology - AC
DevOps Engineer / Docker/REmote
NewApetan Consulting
United States🇺🇸RemoteYesterdayDockerShellAzure+8Technology