Haystack
← Back to Jobs
Technology
CD

Sr InfiniBand AI Network Engineer

Cloud Destinations LLCUnited States🇺🇸United StatesPosted 14 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Sr InfiniBand AI Network Engineer to provide senior InfiniBand engineering services supporting Artificial Intelligence and High Performance Computing clusters where fabric stability and latency sensitive performance are mission critical. This is a hands on role focused on cluster scale deployment, health validation, and advanced troubleshooting of transport, fabric, and endpoint behavior.

Note: This is a W2 only role — C2C, C2H will not be considered

This is a 12 month remote contract opportunity within the United States.Responsibilities:

  • Deploy and validate InfiniBand fabrics supporting distributed Artificial Intelligence training and High Performance Computing workloads.
  • Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior.
  • Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance.
  • Validate switch, host channel adapter, firmware, and cable consistency during cluster deployment and expansion activities.
  • Utilize low level fabric tools to isolate bad links, unstable ports, unhealthy endpoints, topology mismatches, or subnet manager instability.
  • Partner with Linux, deployment, and Artificial Intelligence validation teams to drive root cause analysis from application symptoms to fabric source.
  • Define and execute preflight and post change validation workflows for InfiniBand cluster readiness.
  • Document recurring fault patterns and create remediation playbooks for operational teams. Qualifications:

Required Qualifications

  • 7 or more years of experience in High Performance Computing or advanced networking environments, including direct InfiniBand operations experience.
  • Strong working knowledge of InfiniBand architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics.
  • Experience with RDMA, NCCL related network dependencies, and multi node Artificial Intelligence workload sensitivity to transport issues.
  • Strong troubleshooting skills using fabric health, counter, topology, and endpoint analysis tools.
  • Experience with firmware and driver alignment across host channel adapters, switches, and Linux hosts.
  • Ability to triage complex issues spanning host configuration, fabric state, and application communication behavior.
  • Strong written and verbal communication skills with excellent documentation capabilities. 

Preferred Qualifications

  • Experience supporting DGX, GPU SuperPOD, or equivalent Artificial Intelligence cluster environments.
  • Familiarity with Unified Fabric Manager (UFM), telemetry pipelines, and automated fabric validation.
  • Experience correlating InfiniBand anomalies with Artificial Intelligence training performance outcomes. Tools and Technologies:
  • InfiniBand
  • RDMA
  • NCCL
  • Unified Fabric Manager (UFM)
  • Linux
  • High Performance Computing Networks
  • Artificial Intelligence Infrastructure
  • GPU Clusters
  • Host Channel Adapters
  • Fabric Telemetry
  • Network Diagnostics Tools
  • Performance Analysis Tools
  • Firmware Management

Similar jobs