Quick Overview
Job Description
Job Title: Senior Architect AI Infrastructure & Fabric
Client: LTTS
Location: Plano, TX Hybrid Employment Type: FTE
Experience: 6+ Years HPC or AI Infrastructure Engineering Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service
About the Role
We are building a GPU-as-a-Service and AI Factory practice supporting enterprise and industrial customers. This role focuses on the infrastructure beneath the operating system, including GPU node architecture, compute and storage fabrics, bare-metal provisioning, high-performance storage, and data center infrastructure.
You will be responsible for designing scalable GPU infrastructure and troubleshooting complex cluster, storage, networking, and performance issues.
Key Responsibilities
- Design GPU cluster physical and logical topology, including node configurations, rail-optimized fabric layouts, oversubscription ratios, and failure domains.
- Architect and deploy InfiniBand NDR/XDR and high-performance Ethernet fabrics, including subnet management, UFM, adaptive routing, congestion control, and SHARP offload.
- Design high-performance AI storage architectures using parallel filesystems such as WEKA, Lustre, GPFS, VAST, and DDN.
- Work with GPUDirect Storage, NVMe-oF, and NFS-over-RDMA and size storage for dataloader reads, checkpoint writes, and artifact serving.
- Own bare-metal cluster lifecycle, including provisioning, firmware and driver baselines, imaging, node validation, and burn-in.
- Produce BOMs and infrastructure sizing for compute, networking, optics, and storage.
- Validate infrastructure against customer power, cooling, floor loading, and deployment requirements.
- Run scaling and fabric benchmarks including NCCL bus bandwidth, IB performance tests, IOR, and fio.
- Diagnose interconnect and GPU cluster scaling issues, including link errors, topology binding, NUMA/PCIe affinity, GPUDirect RDMA, and I/O stalls.
- Establish infrastructure standards for offshore delivery teams and review their work before customer delivery.
Required Qualifications
- 6+ years of HPC or AI infrastructure engineering experience.
- Production multi-node GPU cluster experience is mandatory.
- Deep experience with InfiniBand and/or high-performance Ethernet, including fabric design, subnet management, congestion behavior, and troubleshooting.
- Strong experience with NVIDIA GPU platforms, including HGX or DGX-class systems, NVLink, NVSwitch, drivers, and firmware.
- Experience designing and tuning parallel or scale-out filesystems for I/O-intensive workloads.
- Strong bare-metal cluster provisioning and Linux systems engineering experience.
- Understanding of data center infrastructure, including rack power, airflow, direct-liquid cooling concepts, cabling, and optics planning.
Nice to Have
- NVIDIA Enterprise Reference Architecture experience.
- Spectrum-X or BlueField DPU experience.
- Liquid-cooled GPU deployment experience.
- WEKA, VAST, or DDN certification.
- NVIDIA networking certification.
Similar jobs
- VT
Oracle ERP Cloud SCM Functional Consultant (Order Management Lead/SME)
NewVSDL Technologies INC
United States🇺🇸RemoteYesterdayOracleERPRequirements Gathering - IR
Warranty Administrator
NewIngersoll Rand
Quincy, IL🇺🇸$50k - $70k/yrHybridYesterdayMicrosoft OfficeSAP - KT
Remote Sr. QA Playwright / Postman Automation
NewKforce Technology Staffing
Carlsbad, CA🇺🇸RemoteYesterdayAgileComplianceContinuous Improvement+5 - KT
eCommerce Product Insights Senior Analyst
NewKforce Technology Staffing
Miami, FL🇺🇸HybridYesterdaySQLTableauCSAT+7 - GI
Senior Asset Management Consultant / Oracle Fusion Asset Management Consultant
NewGalaxy i Technologies, Inc.
Biscayne Park, FL🇺🇸HybridYesterdayOracle - IS
Manager, Identity & Access Management (IAM) & Automation
NewINSPYR Solutions
Davie, FL🇺🇸$170k - $180k/yrHybridYesterdayNode.jsAWSMFA+14