Quick Overview
Job Description
JOB TITLE: Senior AI Infrastructure Engineer – NVIDIA GPU / GB300
JOB TYPE: Contract
LOCATION: Remote – United States
TRAVEL: Must be willing to travel to the Bay Area / San Jose, CA as required
JOB DESCRIPTION
We are seeking a Senior AI Infrastructure Engineer with strong hands-on experience deploying, configuring, and operating large-scale NVIDIA GPU and HPC infrastructure.
The ideal candidate will have direct experience with NVIDIA GB200, GB300, B200, B300, Blackwell, HGX, NVL72, or comparable GPU platforms. This role requires hands-on experience with rack-scale GPU deployments, cluster provisioning, AI/HPC workload orchestration, high-performance networking, firmware, monitoring, and infrastructure automation.
The successful candidate should be comfortable working across compute, networking, storage, cooling, cluster management, and data center infrastructure.
KEY RESPONSIBILITIES
• Deploy and bring up large-scale NVIDIA GPU clusters and AI infrastructure.
• Perform rack-scale GPU infrastructure integration, provisioning, configuration, validation, and production readiness activities.
• Configure and manage GPU clusters using NVIDIA Base Command Manager or similar cluster management platforms.
• Perform BIOS, BMC, IPMI, Redfish, firmware, and hardware validation activities.
• Support GPU cluster lifecycle management, node provisioning, health monitoring, and troubleshooting.
• Configure and troubleshoot high-performance GPU networking environments including InfiniBand, NVLink, NVSwitch, RoCE, and high-speed Ethernet.
• Support NVIDIA networking platforms, GPU fabrics, and large-scale AI/HPC cluster connectivity.
• Deploy and manage Kubernetes-based GPU environments for AI/ML workloads.
• Support Slurm-based HPC workload scheduling, resource management, queue configuration, and workload optimization.
• Automate infrastructure provisioning, validation, monitoring, and operational workflows using Python, Ansible, Bash, or similar tools.
• Monitor and troubleshoot GPU cluster performance, networking, compute, storage, and infrastructure issues.
• Work with engineering, data center, networking, facilities, and hardware teams to resolve deployment and operational issues.
• Support telemetry, observability, incident response, and infrastructure reliability initiatives.
• Participate in hardware/firmware upgrades, system validation, root-cause analysis, and production support.
REQUIRED SKILLS
• Strong hands-on experience with NVIDIA GPU infrastructure.
• Experience with NVIDIA GB200 / GB300, B200 / B300, Blackwell, HGX, NVL72, or similar GPU platforms.
• Experience with large-scale GPU cluster or AI Factory deployments.
• Experience with GPU provisioning, cluster bring-up, and infrastructure lifecycle management.
• Experience with NVIDIA Base Command Manager or similar HPC cluster management tools.
• Strong Linux administration and troubleshooting skills.
• Experience with Kubernetes and/or Slurm.
• Experience with InfiniBand and/or high-performance Ethernet/RoCE.
• Experience with NVLink / NVSwitch or GPU fabric technologies.
• Experience with firmware, BIOS, BMC, IPMI, Redfish, and hardware validation.
• Strong scripting/automation experience using Python, Bash, Ansible, or similar technologies.
• Experience supporting HPC, AI/ML, or large-scale GPU workloads.
PREFERRED EXPERIENCE
• NVIDIA GB200 / GB300 NVL72
• NVIDIA Blackwell architecture
• DGX SuperPOD
• NVIDIA Spectrum-X
• NVIDIA BlueField DPU
• GPUDirect / RDMA
• CUDA
• Run:ai
• DDN / Lustre / parallel file systems
• Prometheus / Grafana / Ganglia
• Direct-to-chip liquid cooling and GPU data center infrastructure
• Large-scale AI Factory deployments
CANDIDATE PROFILE
We are looking for candidates with demonstrated hands-on experience rather than candidates who have only worked with these technologies at a high level.
Candidates should be able to explain:
• GPU platforms and clusters they have personally deployed or supported
• Number of GPUs, racks, or clusters involved
• Their specific responsibilities during deployment and operations
• Networking technologies used
• Kubernetes / Slurm implementation
• Firmware, provisioning, monitoring, and troubleshooting responsibilities
• AI/HPC workloads supported
LOCATION & TRAVEL
This is a remote position within the United States.
Candidates must be willing and able to travel to the Bay Area / San Jose, California, as required.
Please include the candidate's current location and willingness to travel with each submission.
WORK AUTHORIZATION
SUBMISSION REQUIREMENTS
Please provide the following with each candidate submission:
• Updated Resume
• Current Location
• Work Authorization
• Earliest Availability
• Willingness to Travel to Bay Area / San Jose, CA
• Brief summary of relevant NVIDIA GPU / AI infrastructure experience
IMPORTANT
Please do not submit candidates whose experience is primarily limited to:
• Generic enterprise networking
• Traditional network engineering
• Generic cloud engineering
• Basic Kubernetes administration
• Generic data center operations
• General Linux administration without GPU/HPC experience
• General project management
Candidates with direct NVIDIA GPU, AI Factory, HPC, GB200/GB300, Blackwell, InfiniBand, NVLink/NVSwitch, Kubernetes, Slurm, and large-scale cluster deployment experience will be strongly preferred.
Similar jobs
- SJ
Kubernetes Tech Lead
NewSelby Jennings
Manhattan, NY🇺🇸Hybrid17 hours agoAWSBashJava+3Technology - AT
Security Engineer
NewAivanta Tech Inc
Seattle, WA🇺🇸Hybrid17 hours agoMicroservicesBashLLM+2Technology - GS
Help Desk Analyst (Hybrid Onsite)
NewGSK Solutions Inc.
Trenton, NJ🇺🇸€18/hrHybrid17 hours agoActive DirectoryMicrosoft ExcelTechnology - SY
Junior Java Software Engineer [$248k/yr+] TS/SCI-FS Poly with Security Clearance
NewSYSTOLIC
Annapolis Junction, MD🇺🇸Hybrid17 hours agoDockerMicroservicesEncryption+8Technology - SS
Controls and QA Analyst
NewSouthern Scripts
St. Louis, Missouri🇺🇸Hybrid31 minutes agoHIPAATechnology - MS
Intern - AI Solutions Engineer
NewMidland States Bank
Saint Louis, Missouri🇺🇸Hybrid31 minutes agoSQLMicrosoft ExcelPythonTechnology - AH
Epic Certified EHR Systems Analyst III (EpicCare Ambulatory)
NewAdventist Health
Auburn, California🇺🇸Remote31 minutes agoTechnology - AD
CPI Data Center Project Manager, US-West CPI
NewAmazon Data Services, Inc.
Boardman, Oregon🇺🇸$111.1k/yrHybrid31 minutes agoAWSTechnology - TR
Acquisition Integration Advisor - Salesforce Ecosystem
NewTransunion
Chicago, Illinois🇺🇸$100.1k/yrHybrid31 minutes agoSalesforceCRMTechnology - JM
J.P. Morgan Wealth Management - Senior Product Associate, Service Desktop
NewJ.P. Morgan
Plano, Texas🇺🇸On-site31 minutes agoSalesforceSnowflakeTableau+7Technology - SO
Senior Vulnerability Management Analyst
NewSystem One
Birmingham, Alabama🇺🇸Hybrid31 minutes agoAWSComplianceMicrosoft Excel+2Technology - JM
Senior Lead Infrastructure Engineer (Corporate Structure - Infrastructure Platforms)
NewJ.P. Morgan
Plano, Texas🇺🇸On-site31 minutes agoGCPAWSEncryption+6Technology