Quick Overview
Job Description
hackajob is collaborating with Oracle to connect them with exceptional professionals for this role. Description
Role Overview
The Senior Vice President of Compute Platform will lead the next phase of evolution of the foundational compute platform that underpins the entire Compute and AI Infrastructure portfolio. This executive will be responsible for taking an already established compute platform to the next level of scale, consistency, efficiency, reliability, and architectural maturity.
The role spans the hardware-software interface, standard compute control plane and data plane, host software, lifecycle management, and the core abstractions through which compute infrastructure is provisioned, operated, secured, and consumed. The SVP will drive greater standardization and simplification across the compute stack, strengthen common platform capabilities, improve hardware-software integration, and increase the speed and efficiency with which new compute technologies can be introduced into the fleet.
A critical part of this mandate will be leading the AI transformation of the compute platform and engineering organization. The SVP will identify and drive the use of AI across software development, operations, reliability engineering, fleet management, capacity optimization, incident response, engineering productivity, and decision-making to materially improve how the platform is built and operated. This role requires more than outstanding engineering leadership. The SVP must also operate as a strong business leader with a clear business-outcome mindset.
They will connect technical and AI-driven transformation initiatives to customer impact, growth, cost, efficiency, reliability, margin, and capital productivity, while maintaining deep visibility into the business and operational metrics that define platform performance. Responsibilities
Key Responsibilities
Advance the End-to-End Compute Platform
Own the architecture, strategy, roadmap, execution, and operational outcomes for the common compute platform underpinning the broader compute infrastructure portfolio.
Evolve the platform toward greater consistency, modularity, scalability, reliability, and operational efficiency. Strengthen common platform capabilities, reduce fragmentation, and continuously improve the architecture to support increasing scale and complexity. AI Transformation of the Platform
Lead the transformation of the compute platform and engineering organization through the application of AI. Establish a clear strategy for applying AI across software engineering, platform operations, reliability, capacity management, fleet optimization, testing, observability, and incident management.
Drive adoption of AI-assisted software development and AI-enabled operations to improve productivity, software quality, anomaly detection, root-cause analysis, automated remediation, capacity forecasting, and fleet optimization. Establish measurable goals for engineering productivity, incident reduction, operational efficiency, deployment velocity, utilization, and cost. Hardware-Software Interface
Strengthen the interface between hardware and software engineering to improve how new compute technologies are integrated, qualified, deployed, and operated.
Drive consistent interfaces across firmware, operating systems, device drivers, accelerators, networking, storage, and platform-management components. Compute Control Plane
Own and continuously evolve the standard compute control plane across the compute portfolio. Improve the scale, reliability, efficiency, and consistency of capacity management, scheduling, placement, provisioning, configuration, orchestration, lifecycle management, health management, and fleet operations. Compute Data Plane
Own the evolution of the foundational compute data plane and host software architecture.
Drive greater standardization across virtualization, bare metal, containers, accelerated computing, networking, storage integration, workload isolation, security enforcement, and runtime management. Platform Architecture and Standardization
Drive architectural simplification and standardization across the compute organization. Establish clear and stable contracts between platform layers and higher-level products, strengthen API consistency and backward compatibility, and use AI-assisted analysis and engineering tools to accelerate modernization.
Business Outcome and Operational Leadership
Operate as both an engineering leader and a business leader for the compute platform. Ensure engineering and AI transformation priorities are directly connected to measurable business outcomes including customer experience, revenue growth, infrastructure efficiency, cost reduction, margin improvement, capacity availability, and capital productivity. Reliability, Security, and Operational Excellence
Raise the bar on availability, resiliency, security, performance, and operational efficiency.
Apply AI to accelerate anomaly detection, root-cause analysis, incident response, predictive maintenance, and automated recovery. AI and Accelerated Compute
Evolve the compute platform to support the rapidly increasing scale and complexity of AI and accelerated infrastructure. Strengthen common platform capabilities for GPUs, custom accelerators, large-scale clusters, and emerging AI infrastructure architectures.
Engineering Leadership and Organization
Lead a world-class engineering organization spanning compute platform architecture, systems software, control plane, host software, fleet management, virtualization, and infrastructure services. Build a culture of technical excellence, accountability, simplification, scalability, operational discipline, business ownership, and AI-enabled engineering. Leadership Impact
- Greater consistency and standardization across the compute platform.
- A more scalable, reliable, and efficient compute control plane and data plane.
- Higher fleet reliability, utilization, automation, and operational efficiency.
- Material improvements in engineering productivity through AI-assisted development and operations.
- Greater use of predictive and autonomous capabilities across fleet operations.
- Improved infrastructure economics and capital efficiency.
- A compute platform capable of supporting the next generation of cloud and AI workloads at significantly greater scale.
- 15+ years of experience in large-scale infrastructure, distributed systems, cloud computing, systems software, or related engineering domains.
- Significant senior engineering leadership experience managing large, globally distributed engineering organizations.
- Deep understanding of compute infrastructure and distributed systems.
- Proven experience evolving and operating infrastructure platforms at hyperscale.
- Demonstrated success improving scalability, reliability, operational efficiency, and economics of mission-critical infrastructure.
- Strong track record of driving platform standardization and architectural simplification.
- Demonstrated ability to apply emerging technologies, including AI, to transform engineering productivity and operational effectiveness.
- Strong business and operational acumen, with experience using business, financial, and operational metrics to drive priorities and investments.
Preferred Qualifications
- Experience in a senior leadership role at a leading hyperscale cloud provider.
- Experience owning major portions of a cloud compute platform, host platform, virtualization stack, or infrastructure control plane.
- Deep familiarity with GPU and accelerated computing infrastructure.
- Experience operating infrastructure across hundreds of thousands or millions of servers.
- Experience driving AI-enabled transformation across large software or infrastructure engineering organizations.
- Strong understanding of cloud infrastructure economics, including utilization, cost of capacity, capital efficiency, and unit economics.
This role requires a systems leader who can take a highly capable compute platform and elevate it to the next level of scale, simplicity, standardization, operational excellence, and business impact.
The ideal leader combines deep technical credibility with strong business judgment and a forward-looking view of how AI can transform engineering and infrastructure operations. They will see AI not simply as a workload the platform must support, but as a fundamental tool for transforming how the platform itself is engineered, operated, optimized, and evolved.
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $347,800 to $623,900 per annum. May be eligible for bonus, equity . click apply for full job details
Similar jobs
- IA
UKG Kronos Consultant
NewIT America
Dallas, TX🇺🇸Remote17 hours agoSQLScrumAgile+7 - TU
Consultants
NewTrans Union LLC
Chicago, IL🇺🇸$97.5k - $108k/yrHybrid17 hours agoSQLMachine LearningHadoop+5 - AD
End User Computing
NewAdbakx LLC
Nashville, TN🇺🇸$25/hrOn-site17 hours agoActive DirectoryComplianceContinuous Improvement+3 - EE
Drafter
NewExpress Employment Professionals
Columbus, OH🇺🇸$28 - $37/hrHybrid17 hours agoAutoCADDocument ManagementSAP - RY
AI Platform Architect & Lead Enterprise Data, Analytics
NewRyanBPM
Kansas City, MO🇺🇸On-site17 hours agoLookerBigQueryDatabricks+5 - SB
Oracle Integration Specialist Remote Location
NewSierra Business Solution LLC
United States🇺🇸Remote17 hours agoOracleSQL - AT
Oracle EPM Architect
NewAivanta Tech Inc
United States🇺🇸Remote17 hours agoOracleForecasting - RY
OpenText Exstream Lead
NewRyanBPM
MacArthur, PA🇺🇸On-site17 hours agoTriage - PC
Producer
NewPyramid Consulting, Inc.
New York, NY🇺🇸$42 - $44/hrHybrid17 hours agoAccounts PayableAdobe Creative SuiteCompliance+2 - MA
Fab Technician with Journeyman
NewModern Agile Technologies, LLC
Albany, NY🇺🇸On-site17 hours agoAgileMicrosoft Office - ST
AI Tech Lead
NewSRI Tech Solutions
Charlotte, NC🇺🇸Hybrid17 hours agoDockerMicroservicesAWS+10 - RA
Network Design Engineer
NewRADGOV INC
San Jose, CA🇺🇸On-site17 hours agoDNSHTTPIoT+1