← Back to Jobs
Technology
Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple, Inc.Cupertino, CA🇺🇸United StatesPosted 17 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
People at Apple don't just build products - they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it.
The Apple Service Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and \\"just work.\\" If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you!
Description
Apple Service Engineering (ASE)'s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple's data centers. You will manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies - ensuring platform availability, reliability, and performance at Apple scale.
You will drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation, establishing SLOs, driving production readiness, and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity across all domains.
Minimum Qualifications
5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems
Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes
Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale
Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt)
Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering
Excellent verbal and written communication skills with the ability to influence across teams and levels
Demonstrated ability to recruit, develop, and retain high-performing engineering talent
Preferred Qualifications
Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation
Experience managing or scaling batch compute, job scheduling, or HPC platforms
Proficiency in Go or Python with a strong automation-first mindset
Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale
Experience operating large-scale multi-tenant Infrastructure as a Managed Service
Experience managing geographically distributed teams and follow-the-sun on-call models
Track record of driving capacity efficiency initiatives resulting in measurable cost optimization
The Apple Service Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and \\"just work.\\" If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you!
Description
Apple Service Engineering (ASE)'s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple's data centers. You will manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies - ensuring platform availability, reliability, and performance at Apple scale.
You will drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation, establishing SLOs, driving production readiness, and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity across all domains.
Minimum Qualifications
5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems
Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes
Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale
Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt)
Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering
Excellent verbal and written communication skills with the ability to influence across teams and levels
Demonstrated ability to recruit, develop, and retain high-performing engineering talent
Preferred Qualifications
Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation
Experience managing or scaling batch compute, job scheduling, or HPC platforms
Proficiency in Go or Python with a strong automation-first mindset
Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale
Experience operating large-scale multi-tenant Infrastructure as a Managed Service
Experience managing geographically distributed teams and follow-the-sun on-call models
Track record of driving capacity efficiency initiatives resulting in measurable cost optimization
Skills
OpenStack
Machine Learning
Ansible
Chef
Grafana
Kubernetes
Prometheus
Python
Terraform
Similar jobs
Quality Engineering Manager
Apex Systems · Ruther Glen, United States
3 minutes agoManager, Communication Engineering
Wabtec · Louisville, United States
15 minutes ago$91.1k - $129.8k/yrLead, Digital Engineering
Montefiore Health System Inc · Elmsford, United States
47 minutes ago$156k - $195k/yrPrograms Chief Engineer
Anduril Industries · Costa Mesa, United States
56 minutes ago$254k - $336k/yrSenior Engineering Manager - Incident Response
Atlassian Inc. · United States
59 minutes ago$221.4k - $289.1k/yrSoftware Engineering Manager (TS/SCI)
Vantor · Colorado Springs, United States
1 hour ago$128k - $170k/yr