Quick Overview
Job Description
Position Title: SRE Engineer – Azure, DevOps & Kubernetes
Location: Alpharetta, GA
Years of Exp.: 8-10 Years
Salary: $105,000/Annum
Role Overview:
We are looking for an experienced Site Reliability Engineer with strong hands-on expertise in Microsoft Azure, DevOps engineering, and Kubernetes-based platform operations. The role will focus on improving reliability, scalability, observability, automation, and operational excellence for cloud-native applications and enterprise platforms.
Key Responsibilities:
· Ensure high availability, performance, scalability, and reliability of applications and infrastructure hosted on Azure.
· Manage, operate, and troubleshoot Kubernetes environments, preferably Azure Kubernetes Service (AKS).
· Build and maintain CI/CD pipelines using Azure DevOps, Jenkins, GitHub Actions, or similar tools.
· Implement infrastructure automation using Terraform, ARM templates, Bicep, PowerShell, Bash, or Python.
· Define and track SLIs, SLOs, error budgets, and reliability metrics for production services.
· Set up and manage observability solutions using Azure Monitor, Log Analytics, Application Insights, Prometheus, Grafana, ELK, Splunk, Datadog, or similar tools.
· Participate in incident response, production support, on-call rotation, root cause analysis, and post-incident reviews.
· Develop runbooks, playbooks, automation scripts, and self-healing mechanisms to reduce manual effort and improve MTTR.
· Collaborate with development, DevOps, security, infrastructure, and application teams to improve platform reliability and deployment maturity.
· Implement security best practices across Azure, Kubernetes, CI/CD pipelines, secrets management, access control, and network policies.
Required Skills:
· Strong hands-on experience with Microsoft Azure cloud services, including AKS, Azure Monitor, Azure DevOps, Azure Storage, Azure Networking, Azure Key Vault, Azure AD, and Load Balancers.
· Production experience with Kubernetes, Docker, Helm, ingress controllers, service discovery, scaling, troubleshooting, and cluster operations.
· Good understanding of DevOps practices, CI/CD, Git workflows, branching strategies, release management, deployment automation, and rollback strategies.
· Experience with Infrastructure as Code tools such as Terraform, ARM templates, or Bicep.
· Strong Linux administration, shell scripting, troubleshooting, log analysis, and performance tuning skills.
· Experience with monitoring, logging, alerting, incident management, RCA, and operational dashboards.
· Knowledge of networking concepts including DNS, load balancing, firewalls, SSL/TLS, VNet, subnetting, routing, and Kubernetes networking.
· Experience with ITSM and incident management tools such as ServiceNow, Jira, PagerDuty, or Opsgenie.
· Ability to automate repetitive operational tasks using Python, Bash, PowerShell, or Azure CLI.
· Strong communication skills with the ability to provide clear incident updates, technical documentation, and stakeholder coordination.
Preferred Skills:
· Experience with GitOps tools such as Argo CD or Flux.
· Knowledge of service mesh technologies such as Istio, Linkerd, or Consul.
· Experience with DevSecOps practices, vulnerability scanning, policy enforcement, and container image security.
· Exposure to multi-cloud environments involving AWS or Google Cloud Platform in addition to Azure.
· Experience with chaos engineering, resiliency testing, disaster recovery, and capacity planning.
· Azure certifications such as Azure Administrator, Azure DevOps Engineer Expert, or Azure Solutions Architect are preferred.
Experience & Qualifications:
· Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience.
· 5+ years of experience in SRE, DevOps, Cloud Engineering, Platform Engineering, or Infrastructure Operations.
· 3+ years of hands-on experience with Azure cloud services and production Kubernetes environments.
· Experience supporting enterprise-scale applications in production with defined SLAs/SLOs.
Key Tools & Technologies:
Azure, AKS, Azure DevOps, Kubernetes, Docker, Helm, Terraform, ARM, Bicep, Git, Jenkins, GitHub Actions, Azure Monitor, Log Analytics, Application Insights, Prometheus, Grafana, ELK, Splunk, Datadog, ServiceNow, Jira, PagerDuty, Python, Bash, PowerShell, Azure CLI, Linux.
Candidate Profile:
The ideal candidate should be a technically strong SRE/DevOps engineer who can operate production platforms, automate infrastructure and deployments, troubleshoot complex cloud-native issues, and drive continuous improvement in system reliability, monitoring, and operational maturity.
“Tech Mahindra is an Equal Employment Opportunity employer. We promote and support a diverse workforce at all levels of the company. All qualified applicants will receive consideration for employment without regard to race, religion, color, sex, age, national origin, or disability. All applicants will be evaluated solely on the basis of their ability, competence, and performance of the essential functions of their positions with or without reasonable accommodations. Reasonable accommodations also are available in the hiring process for applicants with disabilities. Candidates can request a reasonable accommodation by contacting the company ADA Coordinator at .”
Similar jobs
- GE
Staff Platform Engineer
NewGemini
New York🇺🇸Remote4 hours agoDockerGCPRust+9Technology - BI
DevOps Engineer with harness and Cloud
NewBURGEON IT SERVICES LLC
Toronto, ON🇺🇸Hybrid20 hours agoDockerShellAWS+4Technology - SC
DevOps Engineer
NewSN Cloud Solutions
Columbus, OH🇺🇸Hybrid20 hours agoDockerMicroservicesShell+27Technology - AC
Devops Engineer with Active Directory
NewAlltech Consulting Services, Inc.
Mountain View, CA🇺🇸Hybrid20 hours agoActive DirectoryAnsibleAzure+8Technology - SE
Staff Site Reliability Engineer - Federal
NewServiceNow
San Diego, CALIFORNIA🇺🇸Hybrid4 hours agoMySQLRubyAWS+4Technology - QB
DevSecOps Engineer
Quantum Business Advisory USA Corp
Washington, DC🇺🇸Hybrid7 weeks agoEngineering