Why This Role Stands Out
This Azure Platform Engineering Manager role offers significant growth opportunities by allowing you to lead critical infrastructure initiatives and leverage cutting-edge Azure services. You'll thrive here if you're a proactive engineer passionate about reliability, automation, and building scalable cloud solutions, and the hybrid work model provides excellent flexibility. Don't miss out on this chance to make a substantial impact with a reputable company.
Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
1 week ago
AzureCapacity PlanningComplianceContinuous ImprovementKubernetesRoot Cause AnalysisTerraformVaultZero Trust
Job Description
Key Responsibilities
Platform Reliability & Operations
- Ensure the availability, performance, scalability, and reliability of Azure-hosted services.
- Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Proactively monitor platform health and performance using observability tooling.
- Perform root cause analysis and implement permanent fixes for recurring incidents.
- Participate in incident management and on-call support rotations where required.
- Lead blameless post-incident reviews, capture lessons learned, and drive corrective actions through to completion.
- Reduce operational toil by identifying repetitive manual tasks and replacing them with automated, reusable engineering solutions.
- Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.
Azure Cloud Engineering
- Design, deploy, and manage Azure infrastructure services including:
- Virtual Networks
- Application Gateways
- API Management
- Azure Kubernetes Service (AKS)
- Azure Firewall
- Azure Storage
- Key Vault
- Azure Monitor
- Azure AI Services
- Implement cloud platform standards and best practices.
- Support multi-region Azure deployments and platform modernisation initiatives.
- Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.
Infrastructure as Code (Terraform)
- Develop and maintain Terraform modules and reusable infrastructure patterns.
- Implement Infrastructure as Code (IaC) standards and governance controls.
- Ensure infrastructure is version controlled, peer-reviewed, and fully automated.
- Manage Terraform state securely and consistently across environments.
DevOps, Automation & AI
- Build and maintain Azure DevOps CI/CD pipelines.
- Automate infrastructure provisioning and application deployments using pipelines with automated delivery and testing routines.
- Implement testing, security scanning, policy compliance, and release gates.
- Support DevOps and platform engineering practices.
- Create automation for operational runbooks, self-healing processes, deployment validation, and environment consistency checks.
- Manage and Implement AI platforms and tools such as Claude & Open AI, to develop skills and support business adoption of agentic AI capabilities.
- Implement and enable self-service approach to technology services.
Resilience, Disaster Recovery & Failover
- Design and implement highly available Azure architectures.
- Develop and maintain disaster recovery and business continuity capabilities.
- Implement and test:
- Regional failover strategies
- Active/Passive architectures
- Active/Active deployments
- Traffic Manager and Front Door failover patterns
- Database resiliency and replication
- Backup and recovery solutions
- Conduct regular resilience and recovery testing exercises.
- Identify and reduce single points of failure across platforms.
- Define and execute game days, chaos testing, and controlled failure scenarios to validate operational resilience.
Security & Governance
- Ensure platforms are secure-by-design.
- Work closely with Security and Architecture teams to implement:
- Zero Trust principles
- RBAC controls / Managed Identities
- Network segmentation and Zone based architecture
- Secrets management
- Support compliance requirements and operational audits.
- Help coordinate security updates, patches, maintenance routines, and upgrades of the underlying system across partners and vendors
- Embed reliability, security, and compliance controls into build and release pipelines to support production readiness.
Continuous Improvement
- Drive automation and reduction of manual operational tasks.
- Improve deployment reliability and platform observability.
- Contribute to architecture standards, runbooks, and operational documentation.
- Partner with engineering, architecture, security, and service teams to define production readiness standards and reliability acceptance criteria.
Similar jobs
- OR
Senior Manager, Core Infrastructure Engineering
NewOracle
Nashville, Tennessee🇺🇸$146.3k - $306.4k/yrHybrid47 minutes agoOracleEncryptionCompliance+4 - LT
Lead, Chief Systems Engineer
NewL3Harris Technologies
Colorado Springs, Colorado🇺🇸$110.5k - $205.5k/yrHybrid1 hour agoScrumAgileKubernetesTechnology - OR
Director, Platform Software Engineering
NewOracle
Santa Clara, California🇺🇸Hybrid2 hours agoOracleTechnology - OR
Director, Core Infrastructure Engineering
NewOracle
Seattle, Washington🇺🇸$121.5k - $306.4k/yrHybrid2 hours agoDockerOracleRust+11 - AS
Technical Manager
NewApex Systems
Sacramento, CA🇺🇸Hybrid15 hours agoAWSETLAzure+1 - GI
Cleared Software Engineering Manager with Security Clearance
NewGA Intelligence
Charlottesville, VA🇺🇸HybridYesterdayScalaAWSMachine Learning+5Technology