Quick Overview
Job Description
Ideal Candidate: A hands-on Senior SRE with expertise in StageProduction deployments, Azure-based microservices, observability (Grafana, Prometheus, Graylog, Azure Monitor), troubleshooting using Application Insights, database operations, and customer escalation management, focused on maintaining highly reliable and available production services.
Position Summary
We are seeking a highly skilled Senior Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational stability of customer-facing applications and services. This role is focused on StageProduction deployments, Monitoring & Observability, Troubleshooting, Incident Management, and Customer Escalation support across Azure-based microservices environments.
The ideal candidate will possess strong experience in cloud operations, microservices, production support, deployment automation, observability platforms, and database troubleshooting, with a proven ability to rapidly diagnose and resolve complex production issues.
Key Responsibilities
Production & Stage Operations
Manage and support Stage and Production environments.
Execute application, infrastructure, configuration, and database deployments.
Validate releases, perform health checks, and coordinate rollback activities.
Support change management and production readiness reviews.
Monitoring & Observability
Build and maintain dashboards, alerts, and monitoring solutions.
Monitor application, infrastructure, and database health using logs, metrics, traces, and telemetry.
Improve observability coverage and reduce alert noise.
Proactively identify reliability and performance issues before customer impact.
Troubleshooting & Incident Response
Troubleshoot software, infrastructure, configuration, deployment, and database-related issues.
Lead incident response activities and production recovery efforts.
Perform root cause analysis (RCA) and implement preventive actions.
Develop operational runbooks and troubleshooting documentation.
Customer Escalation Management
Investigate and resolve customer-reported production issues.
Act as a technical lead during high-priority incidents.
Partner with Engineering, Product, and Customer Support teams to drive issue resolution.
Provide timely communication and status updates during major incidents.
Required Technical Skills
Cloud & Infrastructure
Microsoft Azure
Azure Kubernetes Service (AKS)
Azure Virtual Machines
App Services
Azure Storage
Azure Networking
Application Gateway
Azure Key Vault
Microservices & Containerization
Kubernetes
Docker
Helm Charts
Microservices Architecture
REST APIs
Event-Driven Architecture
Distributed Systems Troubleshooting
CICD & DevOps
Azure DevOps Pipelines
Bitbucket
Git
Helm-based Deployments
CICD Release Management
Deployment Automation
Monitoring & Observability
Grafana
Prometheus
Graylog
Azure Monitor
Application Insights
Log Analytics
Alerting & Dashboard Management
Distributed Tracing
SLISLO Monitoring
Troubleshooting Expertise
Application Performance Issues
Production Incident Management
Configuration & Environment Issues
Deployment Failures & Rollbacks
Kubernetes & Container Troubleshooting
Network & Connectivity Issues
Root Cause Analysis (RCA)
Databases
Azure SQL SQL Server
PostgreSQL MySQL
Cosmos DB
Redis
Query Performance Tuning
Database Monitoring
Backup & Recovery
Automation & Scripting
PowerShell
Python
Bash
Preferred Experience
Supporting enterprise SaaS applications in Production environments.
Azure-based microservices platforms running on AKS.
Customer-facing production support and escalation management.
24x7 on-call and incident response environments.
Site Reliability Engineering (SRE) best practices including SLIs, SLOs, MTTR, and service availability management.
Key Competencies
Strong troubleshooting and analytical skills.
Production support and incident management expertise.
Customer-first mindset.
Excellent communication and stakeholder management.
Ability to perform effectively during critical outages and high-severity incidents.
Continuous improvement and automation mindset.
Similar jobs
- SB
Site Reliability Engineer (SRE) SUNNyVALE - CA - California
NewSierra Business Solution LLC
Sunnyvale, CA🇺🇸Hybrid22 hours agoAzureBashJava+5Technology - BT
Cloud-Native Platform Engineer
NewBraintree Technology Solutions
Dearborn, MI🇺🇸On-site22 hours agoSpringSpring BootAPI Gateway+14Technology - QT
Sr.Devops Engineer
NewQuantom Tech LLC
Atlanta, GA🇺🇸Hybrid22 hours agoDockerScrumAgile+7Technology - BT
Palantir Platform Engineer with Security Clearance
NewBespoke Technologies Inc.
Herndon, VA🇺🇸Hybrid22 hours agoETLPythonTechnology - DS
Azure DevOps Engineer
NewDecision Six Inc.
Southfield, MI🇺🇸Hybrid22 hours agoDockerMicroservicesSSO+6Technology - TS
Senior Observability Engineer
NewTechgene Solutions LLC
United States🇺🇸Remote22 hours agoEngineering