Haystack
← Back to Jobs
Technology
PR

Full-Time - 10+ Principal Site Reliability Engineer - Atlanta, GA (Hybrid)

ProhiresAtlanta, GA🇺🇸United StatesPosted Sep 17, 2026

Quick Overview

Seniority
Leader
Work mode
Hybrid
Location
Atlanta, GA, United States
Posted
18 hours ago
OracleSQLSQL ServerAWSAnsibleAzureDatadogGitHub ActionsGrafanaKubernetesPostgreSQLPowerShellPrometheusPythonTerraform

Job Description

NO H1B'S

 

Principal Site Reliability Engineer
Cloud Engineering | Site Reliability Engineering
Department: Cloud Engineering
Job Type: Full-Time, Individual Contributor
Experience: 10+ Years
Location: Atlanta, GA

Position Summary
We are seeking an experienced Principal Site Reliability Engineer to lead reliability, automation, and operational excellence across our AWS and Azure cloud environments. This role combines deep hands-on technical expertise with strategic vision, with primary responsibility for transforming reactive operational workflows into automated, self-healing systems that scale with our growing SaaS platform.
As a principal-level individual contributor, you will be the subject matter expert for site reliability practices, observability strategy, and operational automation. Your initial focus will be working directly with cases escalated from Support and Professional Services teams, developing deep familiarity with our platforms and common operational patterns. You will then leverage that knowledge to systematically automate the manual steps involved in resolving those cases, driving measurable improvements in operational efficiency and service reliability.

Core Responsibilities
Operational Case Resolution & Automation
● Serve as a primary escalation point for complex infrastructure cases from Support and Professional Services teams across all product lines.
● Diagnose and resolve issues spanning AWS AppStream, RDS, Azure Data Factory, Azure App Services, Kubernetes, and related cloud services.
● Document recurring case patterns, root causes, and resolution steps to build a knowledge base for automation candidates.
● Design and implement automation to eliminate manual intervention in high-frequency case types, targeting measurable reduction in case volume and resolution time.
● Develop self-service tooling and runbooks that enable Support and PS teams to resolve common issues independently.

Site Reliability & Infrastructure Engineering
● Define and drive SLIs, SLOs, and error budgets across production services in both AWS and Azure environments.
● Lead incident response for critical production issues, conduct blameless post-incident reviews, and ensure corrective actions are implemented.
● Design and implement infrastructure improvements that increase system resilience, reduce toil, and improve deployment confidence.
● Partner with engineering teams to improve service architecture, deployment practices, and capacity planning.
● Champion reliability best practices across the organization, including chaos engineering, progressive rollouts, and automated rollback strategies.

Observability & Monitoring
● Own the observability strategy across AWS and Azure, including metrics, logging, tracing, and alerting.
● Design, build, and maintain Grafana dashboards and alerting pipelines that provide actionable visibility into system health and performance.
● Tune monitoring and alerting to reduce noise, improve signal-to-noise ratio, and ensure on-call teams can respond effectively.
● Integrate observability tooling into CI/CD pipelines to enable deployment-aware monitoring.

Automation & Tooling
● Write production-grade automation in Python and PowerShell to address operational toil, infrastructure provisioning, and incident remediation.
● Build and maintain CI/CD pipelines (GitHub Actions, Azure DevOps) for infrastructure and operational automations.
● Create and maintain internal tooling and CLIs that improve operational efficiency for the broader Cloud Operations team.

Technical Leadership & Cross-Functional Collaboration
● Serve as a technical leader and mentor within Cloud Operations, raising the bar on engineering practices and operational rigor.
● Provide architectural guidance on reliability, scalability, and operability for new and existing services.
● Collaborate closely with Development, Support, Professional Services, and Product teams to align reliability efforts with business priorities.
● Lead technical design reviews and contribute to the team’s long-term infrastructure and automation roadmap.

Required Qualifications
● 10+ years of experience in site reliability engineering, platform engineering, DevOps, or cloud infrastructure roles.
● Extensive Windows Server experience.
● Deep hands-on experience with both AWS and Azure, including core compute, networking, storage, and managed services.
● Strong experience with key services such as AWS AppStream, RDS, and Azure Data Factory, App Services, and Kubernetes (AKS/EKS).
● Advanced proficiency in Python and PowerShell for automation, tooling, and scripting.
● Extensive experience building and maintaining CI/CD pipelines using GitHub Actions, Azure DevOps, or similar platforms.
● Strong expertise with observability platforms, particularly Grafana, including dashboard development, alerting configuration, and data source integration.
● Hands-on experience with Terraform or equivalent Infrastructure-as-Code tools for multi-cloud environments.
● Demonstrated ability to analyze operational patterns, identify automation opportunities, and deliver measurable reductions in manual toil.
● Strong understanding of SRE principles including SLIs/SLOs, error budgets, incident management, and blameless postmortems.
● Excellent written and verbal communication skills, with the ability to present technical strategies to both engineering and executive audiences.

Preferred Qualifications
● Relevant certifications such as AWS Solutions Architect Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).
● Experience with additional observability tools such as Prometheus, Loki, Coralogix, or Datadog.
● Background in managing or operating SaaS platforms with multi-tenant architectures.
● Experience with configuration management tools such as Ansible, Automox, or similar.
● Familiarity with database administration for RDS (Oracle, PostgreSQL, SQL Server) and Azure SQL.
● Experience with streaming or data integration platforms such as Azure Data Factory or SnapLogic.
● Track record of building internal developer platforms or self-service operational tooling.
● Background in cloud security practices, including IAM, network segmentation, and secrets management.
What Success Looks Like

First 30 Days
● Complete onboarding to all cloud environments, tooling, and operational workflows across AWS and Azure.
● Begin taking and resolving escalated cases from Support and Professional Services to build platform familiarity.
● Review existing monitoring, alerting, and observability coverage and identify initial gaps.

First 90 Days
● Establish a categorized inventory of recurring case types with documented resolution patterns and automation potential.
● Deliver initial automation for the highest-frequency manual case workflows, demonstrating measurable time savings.
● Present an observability improvement roadmap with prioritized Grafana dashboard and alerting enhancements.
● Begin contributing to CI/CD pipeline improvements and Infrastructure-as-Code standards.

First 6 Months
● Achieve a measurable reduction in manual case volume through automation and self-service tooling.
● Deliver a mature observability layer with actionable dashboards, tuned alerts, and clear SLI/SLO reporting.
● Establish operational automation patterns and reusable tooling that the broader team can adopt and extend.
● Provide technical leadership that measurably elevates the team’s engineering practices and operational maturity.

 

Similar jobs