Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Minneapolis, MN, United States
Posted
19 hours ago
AWSSplunkAzureBashCloudFormationGoogle CloudKubernetesPagerDutyPythonTerraform
Job Description
Title: Azure Cloud DevOps Engineer
Location: Minneapolis, MN
Job Description
Primary Responsibilities:
- Own Reliability Outcomes: Define and track SLIs/SLOs, manage error budgets, and drive continuous improvement to availability, latency, and resiliency for critical services.
- Operate and Mature Observability/AIOps Platform: Build and tune monitoring, dashboards, alerting, and correlation (logs/metrics/traces) using tools like Dynatrace and Splunk to reduce noise and accelerate detection/diagnosis.
- Lead Incident Response & Problem Management: Run on-call/war rooms, conduct root-cause analysis, publish post-incident reviews, and ensure corrective and preventive actions are delivered.
- Automate Toil & Enable Self-Healing: Create runbooks, scripts, and workflows for automated remediation, safe changes, and guardrails to improve MTTR.
- Partner with Engineering & Stakeholders: Consult on architecture, release readiness, capacity planning, and operational standards; translate reliability risks into executive-ready KPI updates (MTTD/MTTR, error budget burn, recurring toil).
Required Skills & Experience:
- AIOps & SRE Fundamentals: 5+ years in SRE/production operations, including SLO/SLI, error budgets, incident management, and automated remediation patterns.
- Observability Toolset: 3+ years building dashboards, alerts, and troubleshooting with tools such as Dynatrace and Splunk (log/metric/trace correlation, alert tuning, noise reduction).
- Cloud & Platform Engineering: 3+ years operating services on AWS/Azure/Google Cloud Platform; strong expertise in Linux, networking, containers/Kubernetes, and IaC (e.g., Terraform/CloudFormation).
- Automation & Agentic Ops Mindset: Proficient in Python/Bash and CI/CD; experienced in building runbooks, self-healing workflows, and participating in on-call rotations.
- AI-Enabled SRE / Intelligent Ops: 1–2+ years applying AI-assisted incident response (auto-summarization, auto-triage, pattern detection) and predictive alerting/anomaly detection in production environments.
- Security / DevSecOps: Working knowledge of vulnerability management, secrets/cert governance, and secure CI/CD gates.
Preferred Skills & Experience:
- Release Safety & Resiliency Engineering: Experience with progressive delivery (canary/blue-green), chaos testing/DR drills, performance/capacity engineering, and well-architected reviews (e.g., Azure WARA/Azure Advisor).
- ITSM & Tooling Integration: Hands-on experience integrating monitoring signals with ServiceNow and PagerDuty to build low-friction escalation workflows.
General Baseline Expectation:
- Enterprise AI Tool Proficiency: Demonstrate consistent use (minimum 90% weekly usage) of enterprise-approved AI tools (e.g., GitHub Copilot, Microsoft 365 Copilot) to enhance coding, documentation, and overall delivery velocity.
Similar jobs
- TC
Site Reliability Engineer (Devops)
NewTECHNEPTUNE CONSULTING INC
Atlanta, GA🇺🇸$55/hrHybrid19 hours agoSQLSQL ServerLoad Balancing+16Technology - XT
Senior Systems Engineer/Senior DevOps Engineer with Security Clearance
NewXTechnologies
San Antonio, TX🇺🇸Hybrid19 hours agoDockerAnsibleBash+5Technology - LE
Principal Software Engineer, DevOps
NewLegion
Remote🇺🇸$220k - $255k/yrRemote2 hours agoDockerMySQLAWS+22Technology - GR
Mid-Level DevSecOps Engineer
NewGRVTY
Herndon🇺🇸Hybrid4 hours agoDockerMicroservicesSQL+22Engineering - KT
Application Operations Engineer
NewKforce Technology Staffing
Jupiter, FL🇺🇸Hybrid19 hours agoSQLAnsibleLESS+1Technology - PA
DevOps Architect - Databricks
NewPantheon
The Woodlands, TX🇺🇸Hybrid19 hours agoAWSAzureDatabricks+1Technology