Haystack
← Back to Jobs
Technology
BT

SRE Engineer

Braintree Technology SolutionsAlpharetta, GA🇺🇸United StatesPosted Sep 22, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Alpharetta, GA, United States
Posted
Yesterday
DockerShellELKLoad BalancingSplunkAzureBashDNSDatadogGitGitHub ActionsGrafanaHelmJenkinsJiraKubernetesPagerDutyPowerShellPrometheusPythonTerraformVault

Job Description

Position Title: SRE Engineer – Azure, DevOps & Kubernetes

Location: Alpharetta, GA

Years of Exp.: 8-10 Years

Key Responsibilities:

  • Ensure high availability, performance, scalability, and reliability of applications and infrastructure hosted on Azure.
  • Manage, operate, and troubleshoot Kubernetes environments, preferably Azure Kubernetes Service (AKS).
  • Build and maintain CI/CD pipelines using Azure DevOps, Jenkins, GitHub Actions, or similar tools.
  • Implement infrastructure automation using Terraform, ARM templates, Bicep, PowerShell, Bash, or Python.
  • Define and track SLIs, SLOs, error budgets, and reliability metrics for production services.
  • Set up and manage observability solutions using Azure Monitor, Log Analytics, Application Insights, Prometheus, Grafana, ELK, Splunk, Datadog, or similar tools.
  • Participate in incident response, production support, on-call rotation, root cause analysis, and post-incident reviews.
  • Develop runbooks, playbooks, automation scripts, and self-healing mechanisms to reduce manual effort and improve MTTR.
  • Collaborate with development, DevOps, security, infrastructure, and application teams to improve platform reliability and deployment maturity.
  • Implement security best practices across Azure, Kubernetes, CI/CD pipelines, secrets management, access control, and network policies.

 Required Skills:

  • Strong hands-on experience with Microsoft Azure cloud services, including AKS, Azure Monitor, Azure DevOps, Azure Storage, Azure Networking, Azure Key Vault, Azure AD, and Load Balancers.
  • Production experience with Kubernetes, Docker, Helm, ingress controllers, service discovery, scaling, troubleshooting, and cluster operations.
  • Good understanding of DevOps practices, CI/CD, Git workflows, branching strategies, release management, deployment automation, and rollback strategies.
  • Experience with Infrastructure as Code tools such as Terraform, ARM templates, or Bicep.
  • Strong Linux administration, shell scripting, troubleshooting, log analysis, and performance tuning skills.
  • Experience with monitoring, logging, alerting, incident management, RCA, and operational dashboards.
  • Knowledge of networking concepts including DNS, load balancing, firewalls, SSL/TLS, VNet, subnetting, routing, and Kubernetes networking.
  • Experience with ITSM and incident management tools such as ServiceNow, Jira, PagerDuty, or Opsgenie.
  • Ability to automate repetitive operational tasks using Python, Bash, PowerShell, or Azure CLI.
  • Strong communication skills with the ability to provide clear incident updates, technical documentation, and stakeholder coordination.

Experience & Qualifications:

  • Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience.
  • 5+ years of experience in SRE, DevOps, Cloud Engineering, Platform Engineering, or Infrastructure Operations.
  • 3+ years of hands-on experience with Azure cloud services and production Kubernetes environments.
  • Experience supporting enterprise-scale applications in production with defined SLAs/SLOs.

Similar jobs