Quick Overview
Seniority
Leader
Work mode
Remote
Location
San Francisco, CA, United States
Posted
Yesterday
OracleCDNKubernetesPythonScheduling
Job Description
Head of AI Infrastructure
United States
$250-350K + Bonus + Equity
/authorized to work in the Canada
Remote with travel
About the company
Our client is a venture-backed energy technology company building a new kind of AI compute network. Instead of one large data center, they deploy liquid-cooled NVIDIA GPU inference pods across a nationwide network of existing sites, running on power that is already permitted and in place.
The pods are designed to be modular and quick to deploy. They avoid new utility permits, water hookups, backup generators and major construction. Each pod shares its site's electrical service with other energy infrastructure, so the platform balances power intelligently and fails over gracefully between sites.
The role
This is the company's first dedicated infrastructure hire, reporting directly to the CTO.
You'll start hands-on and own the AI platform end to end: compute, memory, network, storage, orchestration and field operations. As the fleet grows, you'll build and lead the infrastructure team. You'll also be the senior technical voice for customers and set the 12 to 18 month infrastructure roadmap.
The work moves fast. New NVIDIA drivers ship weekly, new vLLM versions every two weeks and new models monthly, so continuous benchmarking and safe rollouts are at the heart of the job.
What you'll do
What you'll bring
Nice to have
You don't need to check every box to be considered.
Location, compensation and how to apply
United States
$250-350K + Bonus + Equity
/authorized to work in the Canada
Remote with travel
About the company
Our client is a venture-backed energy technology company building a new kind of AI compute network. Instead of one large data center, they deploy liquid-cooled NVIDIA GPU inference pods across a nationwide network of existing sites, running on power that is already permitted and in place.
The pods are designed to be modular and quick to deploy. They avoid new utility permits, water hookups, backup generators and major construction. Each pod shares its site's electrical service with other energy infrastructure, so the platform balances power intelligently and fails over gracefully between sites.
The role
This is the company's first dedicated infrastructure hire, reporting directly to the CTO.
You'll start hands-on and own the AI platform end to end: compute, memory, network, storage, orchestration and field operations. As the fleet grows, you'll build and lead the infrastructure team. You'll also be the senior technical voice for customers and set the 12 to 18 month infrastructure roadmap.
The work moves fast. New NVIDIA drivers ship weekly, new vLLM versions every two weeks and new models monthly, so continuous benchmarking and safe rollouts are at the heart of the job.
What you'll do
- Bring up, burn in and run GPU pods in production, and own the acceptance benchmarks every new pod and hardware generation must pass
- Measure and improve inference performance across compute, memory, network and storage: tokens per second per kW, KV-cache offload, RoCEv2 and NCCL fabric performance, model cold-start
- Test and roll out frequent driver, vLLM and model updates safely, with clear benchmarks and rollback plans
- Run power-aware operations: GPU power caps, curtailment and workload drain coordinated with other on-site energy loads
- Design for graceful failure, so work hands off cleanly between pods and sites
- Operate a multi-tenant, bare-metal Kubernetes GPU platform against service level objectives, with 24/7 incident response
- Write the runbooks field technicians follow at unmanned sites
- Act as technical lead for customers, and set the 12 to 18 month infrastructure roadmap
- Hire, build and lead the infrastructure team as the fleet scales
What you'll bring
- 12+ years in infrastructure, with hands-on ownership of physical production compute at scale (not only consuming public cloud services)
- Experience at a hyperscaler, GPU cloud or neocloud: you know how large organizations keep large systems running
- Practical AI inference knowledge, such as vLLM or other model serving frameworks
- Depth in at least two of: inference serving, InfiniBand or RoCE fabrics, distributed storage, bare-metal Kubernetes
- Experience running distributed or multi-site infrastructure
- Strong Linux skills, plus Python or Go: you read the code and write the fix
- Technical leadership experience, and the ambition to build and manage a team
Nice to have
- Edge or distributed-site infrastructure, such as CDN points of presence, cloud local zones or telecom edge
- Experience with 1,000+ GPU clusters
- Power-aware scheduling, demand response or GPU cluster power management
- Background at a GPU, chip or infrastructure vendor, such as NVIDIA, AMD, Intel or Oracle
You don't need to check every box to be considered.
Location, compensation and how to apply
- Location: fully remote, anywhere in the US or Canada
- Travel: to pod sites and customers as needed, plus two company-wide gatherings a year
- Work authorization: must already be authorized to work in the US or Canada; visa sponsorship is not available
- Compensation: competitive base salary with significant equity
Similar jobs
- OR
Oracle Health - Lead Healthcare Executive (Cerner Millennium)
NewOracle
United States🇺🇸$97.5k - $209.5k/yrHybrid5 minutes agoOracleAccount ManagementBusiness Development+1 - DG
Commissioning Technician - Power Systems
NewD2B Groups
Raleigh, North Carolina🇺🇸Hybrid6 hours ago - NW
UI Artist (Game Industry)
NewNimble Workforce
North Carolina🇺🇸Remote1 hour agoFigmaAdobe Creative SuiteCSS+5 - NW
Technical UI Artist (Game Industry)
NewNimble Workforce
North Carolina🇺🇸Remote1 hour agoFigmaAdobe Creative SuiteCSS+3 - JC
Digital Media Activation Manager
NewJeffreyM Consulting
United States🇺🇸Remote5 hours agoDigital MarketingGoogle AdsSEM+1 - SY
Integration Analyst - Healthcare
NewSymmetrio
United States🇺🇸Remote6 hours agoDockerAWSCerner+8