Haystack
← Back to Jobs
Remote
Technology
TA

Senior AI Infrastructure Engineer – NVIDIA GPU / GB300

Texnere Americas IncUnited States🇺🇸United StatesPosted Oct 8, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
17 hours ago
AnsibleBashCUDAGrafanaKubernetesPrometheusPython

Job Description

JOB TITLE: Senior AI Infrastructure Engineer – NVIDIA GPU / GB300

JOB TYPE: Contract

LOCATION: Remote – United States

TRAVEL: Must be willing to travel to the Bay Area / San Jose, CA as required

JOB DESCRIPTION

We are seeking a Senior AI Infrastructure Engineer with strong hands-on experience deploying, configuring, and operating large-scale NVIDIA GPU and HPC infrastructure.

The ideal candidate will have direct experience with NVIDIA GB200, GB300, B200, B300, Blackwell, HGX, NVL72, or comparable GPU platforms. This role requires hands-on experience with rack-scale GPU deployments, cluster provisioning, AI/HPC workload orchestration, high-performance networking, firmware, monitoring, and infrastructure automation.

The successful candidate should be comfortable working across compute, networking, storage, cooling, cluster management, and data center infrastructure.

KEY RESPONSIBILITIES

• Deploy and bring up large-scale NVIDIA GPU clusters and AI infrastructure.

• Perform rack-scale GPU infrastructure integration, provisioning, configuration, validation, and production readiness activities.

• Configure and manage GPU clusters using NVIDIA Base Command Manager or similar cluster management platforms.

• Perform BIOS, BMC, IPMI, Redfish, firmware, and hardware validation activities.

• Support GPU cluster lifecycle management, node provisioning, health monitoring, and troubleshooting.

• Configure and troubleshoot high-performance GPU networking environments including InfiniBand, NVLink, NVSwitch, RoCE, and high-speed Ethernet.

• Support NVIDIA networking platforms, GPU fabrics, and large-scale AI/HPC cluster connectivity.

• Deploy and manage Kubernetes-based GPU environments for AI/ML workloads.

• Support Slurm-based HPC workload scheduling, resource management, queue configuration, and workload optimization.

• Automate infrastructure provisioning, validation, monitoring, and operational workflows using Python, Ansible, Bash, or similar tools.

• Monitor and troubleshoot GPU cluster performance, networking, compute, storage, and infrastructure issues.

• Work with engineering, data center, networking, facilities, and hardware teams to resolve deployment and operational issues.

• Support telemetry, observability, incident response, and infrastructure reliability initiatives.

• Participate in hardware/firmware upgrades, system validation, root-cause analysis, and production support.

REQUIRED SKILLS

• Strong hands-on experience with NVIDIA GPU infrastructure.

• Experience with NVIDIA GB200 / GB300, B200 / B300, Blackwell, HGX, NVL72, or similar GPU platforms.

• Experience with large-scale GPU cluster or AI Factory deployments.

• Experience with GPU provisioning, cluster bring-up, and infrastructure lifecycle management.

• Experience with NVIDIA Base Command Manager or similar HPC cluster management tools.

• Strong Linux administration and troubleshooting skills.

• Experience with Kubernetes and/or Slurm.

• Experience with InfiniBand and/or high-performance Ethernet/RoCE.

• Experience with NVLink / NVSwitch or GPU fabric technologies.

• Experience with firmware, BIOS, BMC, IPMI, Redfish, and hardware validation.

• Strong scripting/automation experience using Python, Bash, Ansible, or similar technologies.

• Experience supporting HPC, AI/ML, or large-scale GPU workloads.

PREFERRED EXPERIENCE

• NVIDIA GB200 / GB300 NVL72

• NVIDIA Blackwell architecture

• DGX SuperPOD

• NVIDIA Spectrum-X

• NVIDIA BlueField DPU

• GPUDirect / RDMA

• CUDA

• Run:ai

• DDN / Lustre / parallel file systems

• Prometheus / Grafana / Ganglia

• Direct-to-chip liquid cooling and GPU data center infrastructure

• Large-scale AI Factory deployments

CANDIDATE PROFILE

We are looking for candidates with demonstrated hands-on experience rather than candidates who have only worked with these technologies at a high level.

Candidates should be able to explain:

• GPU platforms and clusters they have personally deployed or supported

• Number of GPUs, racks, or clusters involved

• Their specific responsibilities during deployment and operations

• Networking technologies used

• Kubernetes / Slurm implementation

• Firmware, provisioning, monitoring, and troubleshooting responsibilities

• AI/HPC workloads supported

LOCATION & TRAVEL

This is a remote position within the United States.

Candidates must be willing and able to travel to the Bay Area / San Jose, California, as required.

Please include the candidate's current location and willingness to travel with each submission.

WORK AUTHORIZATION

SUBMISSION REQUIREMENTS

Please provide the following with each candidate submission:

• Updated Resume

• Current Location

• Work Authorization

• Earliest Availability

• Willingness to Travel to Bay Area / San Jose, CA

• Brief summary of relevant NVIDIA GPU / AI infrastructure experience

IMPORTANT

Please do not submit candidates whose experience is primarily limited to:

• Generic enterprise networking

• Traditional network engineering

• Generic cloud engineering

• Basic Kubernetes administration

• Generic data center operations

• General Linux administration without GPU/HPC experience

• General project management

Candidates with direct NVIDIA GPU, AI Factory, HPC, GB200/GB300, Blackwell, InfiniBand, NVLink/NVSwitch, Kubernetes, Slurm, and large-scale cluster deployment experience will be strongly preferred.

Similar jobs