Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
Atlanta, GA, United States
Posted
Yesterday
SQLSQL ServerLoad BalancingSplunkTCP/IPAzureBashCDNCloudflareDNSGoogle CloudHTTPKubernetesPowerShellPythonRabbitMQStakeholder ManagementTerraformWAF
Job Description
Site Reliability Engineer
Location: Midtown Atlanta
Duration: Fulltime Permanent
Type: Hybrid - 3 Days onsite
Role Summary
Lead the reliability, scalability, security, and operational excellence of customer-facing platforms across Azure, Google Cloud Platform, and Kubernetes environments. Drive production stability through automation, observability, incident management, and continuous improvement initiatives.
Key Responsibilities
- Lead platform reliability, availability, and performance initiatives.
- Design and support cloud infrastructure in Azure and Google Cloud Platform.
- Manage and optimize Kubernetes environments and containerized applications.
- Implement observability and monitoring using Splunk, AppDynamics, and cloud-native tools.
- Support Cloudflare, Zscaler, SQL Server, RabbitMQ, and enterprise networking components.
- Lead major incident response, RCA, and problem management activities.
- Develop automation and self-healing solutions to improve operational efficiency.
- Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance customer experience and platform stability.
- Serve as a technical escalation point for critical production and customer issues.
Required Skills
- 5+ years of experience in SRE, DevOps, Cloud Operations, or Infrastructure Engineering.
- Strong expertise in Azure, Google Cloud Platform, Kubernetes, Cloudflare, Splunk, AppDynamics, SQL Server, RabbitMQ, and Zscaler.
- Solid networking knowledge (DNS, TCP/IP, HTTP/S, CDN, WAF, Load Balancing, SSL/TLS, Firewalls, VPNs).
- Experience with automation and scripting (Python, PowerShell, Bash, Terraform).
- Strong customer-facing communication and stakeholder management skills.
Core Principles
- Automation First – Eliminate manual effort through automation and self-healing systems.
- Observability-Driven Operations – Leverage logs, metrics, traces, and analytics to proactively identify and resolve issues.
- AI-Powered Reliability – Utilize AI and operational intelligence to accelerate detection, diagnosis, and remediation.
- Customer-Centric Mindset – Prioritize customer experience, stability, and business outcomes.
- Operational Excellence – Continuously improve reliability, scalability, and security.
Success Measures
- Service Availability & Uptime
- SLA/SLO Compliance
- MTTR Reduction
- Incident Reduction
- Automation Adoption
- Customer Satisfaction (CSAT)
- Platform Performance & Stability Improvements
Similar jobs
- QT
Senior Cloud Platform Engineer
NewQUANTUM TECHNOLOGIES LLC
Dallas, TX🇺🇸$85/hrHybridYesterdaySQLAzureDNS+3Technology - BI
Cloud Platform Engineer
NewBMR Infotek
Newark, NJ🇺🇸HybridYesterdayMLOpsAzureGoogle Cloud+5Technology - KR
Devops Engineer
K-Tek Resourcing LLC
Irving, TX🇺🇸Hybrid2 months agoDockerAWSSonarQube+13Technology - RD
SRE (Site Reliability Engineering)
NewRandstad Digital
Hollywood, FL🇺🇸$70 - $80/hrHybridYesterdayAWSCDNTerraformTechnology - CG
Site Reliability Engineer
NewCharter Global, Inc.
Atlanta, GA🇺🇸On-siteYesterdaySQLSQL ServerLoad Balancing+16Technology - TD
DevOps Engineer with Security Clearance
NewTetrad Digital Integrity (TDI)
Wash, DC🇺🇸$180k - $190k/yrOn-siteYesterdayDockerSQLAWS+14Technology