Haystack
← Back to Jobs
Other
VC

Senior Architect AI Infrastructure & Fabric

VST Consulting, IncPlano, TX🇺🇸United StatesPosted Sep 17, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Plano, TX, United States
Posted
Yesterday
Airflow

Job Description

Job Title: Senior Architect AI Infrastructure & Fabric
Client: LTTS
Location: Plano, TX Hybrid Employment Type: FTE
Experience: 6+ Years HPC or AI Infrastructure Engineering Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service

About the Role

We are building a GPU-as-a-Service and AI Factory practice supporting enterprise and industrial customers. This role focuses on the infrastructure beneath the operating system, including GPU node architecture, compute and storage fabrics, bare-metal provisioning, high-performance storage, and data center infrastructure.

You will be responsible for designing scalable GPU infrastructure and troubleshooting complex cluster, storage, networking, and performance issues.

Key Responsibilities

  • Design GPU cluster physical and logical topology, including node configurations, rail-optimized fabric layouts, oversubscription ratios, and failure domains.
  • Architect and deploy InfiniBand NDR/XDR and high-performance Ethernet fabrics, including subnet management, UFM, adaptive routing, congestion control, and SHARP offload.
  • Design high-performance AI storage architectures using parallel filesystems such as WEKA, Lustre, GPFS, VAST, and DDN.
  • Work with GPUDirect Storage, NVMe-oF, and NFS-over-RDMA and size storage for dataloader reads, checkpoint writes, and artifact serving.
  • Own bare-metal cluster lifecycle, including provisioning, firmware and driver baselines, imaging, node validation, and burn-in.
  • Produce BOMs and infrastructure sizing for compute, networking, optics, and storage.
  • Validate infrastructure against customer power, cooling, floor loading, and deployment requirements.
  • Run scaling and fabric benchmarks including NCCL bus bandwidth, IB performance tests, IOR, and fio.
  • Diagnose interconnect and GPU cluster scaling issues, including link errors, topology binding, NUMA/PCIe affinity, GPUDirect RDMA, and I/O stalls.
  • Establish infrastructure standards for offshore delivery teams and review their work before customer delivery.

Required Qualifications

  • 6+ years of HPC or AI infrastructure engineering experience.
  • Production multi-node GPU cluster experience is mandatory.
  • Deep experience with InfiniBand and/or high-performance Ethernet, including fabric design, subnet management, congestion behavior, and troubleshooting.
  • Strong experience with NVIDIA GPU platforms, including HGX or DGX-class systems, NVLink, NVSwitch, drivers, and firmware.
  • Experience designing and tuning parallel or scale-out filesystems for I/O-intensive workloads.
  • Strong bare-metal cluster provisioning and Linux systems engineering experience.
  • Understanding of data center infrastructure, including rack power, airflow, direct-liquid cooling concepts, cabling, and optics planning.

Nice to Have

  • NVIDIA Enterprise Reference Architecture experience.
  • Spectrum-X or BlueField DPU experience.
  • Liquid-cooled GPU deployment experience.
  • WEKA, VAST, or DDN certification.
  • NVIDIA networking certification.

Similar jobs