Sr InfiniBand AI Network Engineer
Quick Overview
Job Description
Sr InfiniBand AI Network Engineer to provide senior InfiniBand engineering services supporting Artificial Intelligence and High Performance Computing clusters where fabric stability and latency sensitive performance are mission critical. This is a hands on role focused on cluster scale deployment, health validation, and advanced troubleshooting of transport, fabric, and endpoint behavior.
This is a 12 month remote contract opportunity within the United States.Responsibilities:
- Deploy and validate InfiniBand fabrics supporting distributed Artificial Intelligence training and High Performance Computing workloads.
- Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior.
- Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance.
- Validate switch, host channel adapter, firmware, and cable consistency during cluster deployment and expansion activities.
- Utilize low level fabric tools to isolate bad links, unstable ports, unhealthy endpoints, topology mismatches, or subnet manager instability.
- Partner with Linux, deployment, and Artificial Intelligence validation teams to drive root cause analysis from application symptoms to fabric source.
- Define and execute preflight and post change validation workflows for InfiniBand cluster readiness.
- Document recurring fault patterns and create remediation playbooks for operational teams. Qualifications:
Required Qualifications
- 7 or more years of experience in High Performance Computing or advanced networking environments, including direct InfiniBand operations experience.
- Strong working knowledge of InfiniBand architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics.
- Experience with RDMA, NCCL related network dependencies, and multi node Artificial Intelligence workload sensitivity to transport issues.
- Strong troubleshooting skills using fabric health, counter, topology, and endpoint analysis tools.
- Experience with firmware and driver alignment across host channel adapters, switches, and Linux hosts.
- Ability to triage complex issues spanning host configuration, fabric state, and application communication behavior.
- Strong written and verbal communication skills with excellent documentation capabilities.
Preferred Qualifications
- Experience supporting DGX, GPU SuperPOD, or equivalent Artificial Intelligence cluster environments.
- Familiarity with Unified Fabric Manager (UFM), telemetry pipelines, and automated fabric validation.
- Experience correlating InfiniBand anomalies with Artificial Intelligence training performance outcomes. Tools and Technologies:
- InfiniBand
- RDMA
- NCCL
- Unified Fabric Manager (UFM)
- Linux
- High Performance Computing Networks
- Artificial Intelligence Infrastructure
- GPU Clusters
- Host Channel Adapters
- Fabric Telemetry
- Network Diagnostics Tools
- Performance Analysis Tools
- Firmware Management
Similar jobs
Technical Support Specialist I
Infobahn Softworld Inc. · Dallas, United States
Just nowSystems Administrator with Security Clearance
Saalex · Ridgecrest, United States
Just now$75k - $82k/yrLead GPS Ground Control Systems Engineer with Security Clearance
MITRE Corporation · Colorado Springs, United States
Just now$146.4k - $183k/yrSalesforce Architect with Contact Center transformation exp
Rivago infotech inc · United States
Just nowLead IAM Security Engineer - Saviynt IGA Or Cyber Ark PAM
Econosoft · Lutz, United States
Just nowMachine Learning Engineer - Graph ML & Code Intelligence
Comtech Global · United States
Just now