Required Staff Infrastructure Engineer – GPU / AI Infrastructure - San Francisco - Onsite
Quick Overview
Job Description
Staff Infrastructure Engineer – GPU / AI Infrastructure - San Francisco - Onsite
Experience Level: Senior / Staff
Industry: AI Infrastructure / Cloud / Data Center
Primary Focus: GPU Infrastructure, Linux, Kubernetes, Bare-Metal Automation
Job Overview
We are seeking a highly experienced Staff Infrastructure Engineer to help build, operate, automate, and scale cloud and bare-metal infrastructure supporting large-scale GPU and AI workloads.
The ideal candidate has strong hands-on experience with GPU hardware, Linux, enterprise servers, bare-metal provisioning, Kubernetes, and infrastructure automation. You will play a critical role in operating the company's infrastructure environment, troubleshooting complex hardware issues, developing automation, and scaling infrastructure to support large deployments.
GPU infrastructure experience is a key requirement for this position.
This role is well suited for an infrastructure engineer who enjoys working across hardware, operating systems, cloud infrastructure, Kubernetes, and automation and is comfortable troubleshooting problems from the physical server layer through the software stack.
What You'll Do
- Manage and maintain the day-to-day operations of cloud and bare-metal infrastructure.
- Support and operate large-scale GPU infrastructure used for AI/ML and high-performance computing workloads.
- Troubleshoot server and hardware issues, with a strong focus on GPUs, GPU connectivity, performance, and reliability.
- Diagnose hardware, firmware, BIOS, BMC/IPMI, networking, and operating system issues.
- Work directly with hardware vendors and data center partners to troubleshoot and resolve hardware failures.
- Develop automation tools and workflows to streamline server provisioning, configuration, and deployment.
- Improve provisioning processes to reduce deployment and SLA turnaround times.
- Design and implement infrastructure capable of supporting mass server deployments, including approximately 80–100 servers simultaneously.
- Perform and automate bare-metal server provisioning, including PXE-based deployment workflows.
- Manage enterprise server technologies including BMC/IPMI, BIOS, firmware, and remote management interfaces.
- Help transition infrastructure toward Kubernetes and containerized workloads.
- Deploy, configure, maintain, and troubleshoot Kubernetes clusters and associated infrastructure.
- Support containerized applications and workloads using technologies such as Docker.
- Develop infrastructure automation and configuration management using tools such as Ansible and Terraform.
- Build tooling and automation using Python and/or Golang.
- Identify recurring infrastructure problems and develop scalable, automated solutions.
- Collaborate with infrastructure, platform, software, data center, and operations teams to improve reliability and scalability.
- Participate in incident response, root-cause analysis, capacity planning, and infrastructure optimization.
- Create documentation, operational procedures, and repeatable deployment processes.
Required Qualifications
- Strong hands-on experience with GPU infrastructure and GPU troubleshooting.
- Significant experience managing Linux-based infrastructure, preferably in large-scale production environments.
- Strong understanding of enterprise-grade server hardware and infrastructure.
- Experience troubleshooting physical servers, including:
- GPUs
- CPUs
- Memory
- Storage
- Network interfaces
- PCIe devices
- Power-related hardware issues
- Experience with bare-metal server provisioning and deployment.
- Hands-on experience with PXE booting and automated server provisioning.
- Experience with BMC/IPMI, BIOS configuration, firmware, and remote server management.
- Experience managing infrastructure at scale and supporting large numbers of servers.
- Hands-on experience with Kubernetes, either as a Kubernetes administrator or developer.
- Experience with containerization technologies, preferably Docker.
- Strong troubleshooting, problem-solving, and systems engineering skills.
- Ability to work across hardware, operating systems, networking, and software infrastructure layers.
Preferred Qualifications
- Experience operating GPU clusters or AI/ML infrastructure at scale.
- Experience with NVIDIA GPUs, GPU servers, or high-density compute infrastructure.
- Experience deploying and managing Kubernetes clusters in production.
- Experience with MAAS (Metal as a Service) or similar bare-metal provisioning platforms.
- Proficiency in Golang or Python for infrastructure automation and tooling.
- Experience with Ansible for configuration management and automation.
- Experience with Terraform and Infrastructure as Code (IaC).
- Experience with Linux system administration and performance troubleshooting.
- Experience with automated OS deployment and infrastructure lifecycle management.
- Experience working with data center hardware and remote hands/vendor teams.
- Familiarity with high-performance computing, distributed systems, or AI/ML infrastructure.
Technical Skills
Required / Core:
- GPU Infrastructure
- GPU Troubleshooting
- Linux
- Bare-Metal Servers
- PXE Boot
- Server Provisioning
- BMC / IPMI
- BIOS / Firmware
- Enterprise Server Hardware
- Kubernetes
- Docker / Containers
Preferred:
- MAAS
- Python
- Golang
- Ansible
- Terraform
- Infrastructure as Code
- Kubernetes Administration
- Kubernetes Deployment
- Automation
- NVIDIA GPU Infrastructure
- AI/ML Infrastructure
Key Responsibilities at a Glance
- GPU Infrastructure: Operate and troubleshoot GPU servers and AI infrastructure.
- Linux: Manage and troubleshoot production Linux environments.
- Bare Metal: Automate PXE-based provisioning and server deployment.
- Scale: Support simultaneous deployments of approximately 80–100 servers.
- Kubernetes: Help migrate infrastructure to Kubernetes and containerized workflows.
- Automation: Build tools and workflows to improve provisioning and operational efficiency.
- Hardware: Diagnose server, GPU, BIOS, BMC/IPMI, firmware, and component-level issues.
- Vendor Management: Work with hardware vendors to resolve complex infrastructure issues.
Ideal Candidate Profile
The ideal candidate is a hands-on Staff Infrastructure Engineer / Senior Infrastructure Engineer / Linux Infrastructure Engineer / GPU Infrastructure Engineer / Systems Engineer with deep experience operating physical infrastructure at scale.
You should be comfortable moving between GPU hardware troubleshooting, Linux administration, bare-metal provisioning, Kubernetes, and infrastructure automation. You enjoy solving difficult infrastructure problems, eliminating repetitive operational work through automation, and building systems that can reliably scale to large numbers of servers.
GPU infrastructure and hardware troubleshooting experience are especially important for this role.
Similar jobs
- IS
EPIC Application Analyst - ADT / Prelude - Epic OpTime Certified
Newi4 Search Group Healthcare
White Plains, New York🇺🇸$91.0k - $136.5k/yrHybrid15 minutes agoPatient CareRequirements GatheringWorkdayTechnology - PF
Application Specialist - Taylor, MO
NewPrairieland FS
Taylor, Missouri🇺🇸Hybrid15 minutes agoComplianceInventory ManagementTechnology - FL
LIS Application Specialist
NewFrontage Laboratories
Exton, Pennsylvania🇺🇸Hybrid15 minutes agoGCPClinical TrialsGDPR+1Technology - BR
Application Specialist
NewBraskem
Philadelphia, Pennsylvania🇺🇸Hybrid15 minutes agoSQLSQL ServerAzure+3Technology - SP
Software Engineer IV – ERP Finance (Oracle/PLSQL)
NewSPECTRUM
Kannapolis, NC🇺🇸Hybrid23 hours agoOraclePL/SQLSQL+6Technology - SP
Software Engineer IV – ERP Finance (Oracle/PLSQL)
NewSPECTRUM
SAINT LOUIS, MO🇺🇸Hybrid23 hours agoOraclePL/SQLSQL+6Technology