Haystack
← Back to Jobs
Other
VC

Principal Architect GPU Platform & Orchestration

VST Consulting, IncPlano, TX🇺🇸United StatesPosted Sep 17, 2026

Quick Overview

Seniority
Leader
Work mode
Hybrid
Location
Plano, TX, United States
Posted
Yesterday
AnsibleHelmKubernetesOnboardingSchedulingTerraform

Job Description

Job Title: Principal Architect GPU Platform & Orchestration
Client: LTTS
Location: Plano, TX Hybrid Employment Type: FTE
Rate: $/hr on 1099 Experience: 7+ Years Platform Engineering; 4+ Years Kubernetes in Production Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service

About the Role
We are building a GPU-as-a-Service and AI factory practice from the ground up, delivering multi
tenant GPU platforms for enterprise and industrial customers. This is the senior technical seat on
that platform. You will own the orchestration and multi-tenancy architecture that turns a GPU
cluster into a consumable service, set the standards our global delivery team builds against, and
serve as deputy to the practice lead in customer architecture engagements.
This is an architecture role. You will design, review, and defend and lead an offshore
engineering pod that executes.
What you ll do

Architect the GPUaaS control plane on Kubernetes and OpenShift: NVIDIA GPU Operator,
Network Operator, device plugin, MIG manager, node feature discovery.
Design multi-tenancy end to end MIG partitioning strategy, time-slicing tiers, namespace
and RBAC model, network policy, quotas, priority classes, and tenant onboarding.
Own GPU scheduling and allocation policy: gang scheduling (Kueue, Volcano), fair-share
and preemption, topology-aware placement, and Slurm integration where customers run
genuine batch HPC.
Define the service catalog instance shapes, self-service request flow, and GPU metering
for chargeback or showback from DCGM telemetry.
Build and own the reusable platform blueprint: reference architecture, Terraform and
Helm modules, GitOps patterns, and runbooks that every engagement starts from.
Technically lead an offshore delivery pod set standards, run design reviews, gate
deliverables before they reach a customer.
Partner with the practice lead on customer discovery, solution design, and technical
escalation; lead design sessions independently as the practice scales.
What you need
7+ years platform engineering, with 4+ on Kubernetes in production; OpenShift experience
valued.
Demonstrated GPU workload orchestration GPU Operator, MIG, device plugin, GPU
scheduling policy on real multi-node clusters.
Real multi-tenancy design experience: isolation, quota, RBAC, network segmentation, and
the failure modes each produces.
Batch or HPC scheduling background (Slurm, LSF, PBS) or gang scheduling on Kubernetes.
Strong IaC and GitOps: Terraform, Helm, Argo CD or Flux, Ansible.
Experience leading distributed or offshore engineering teams through written standards
rather than direct supervision.
Customer-facing credibility you can whiteboard a design for a CTO and defend it under
challenge.
Nice to have
NVIDIA AI Enterprise; Run:ai or equivalent GPU orchestration; consulting or professional-services
background; internal developer platform / service catalog experience; CKA or CKS.
First 90 days
v1 GPUaaS platform blueprint published and deployed in our reference environment; tenancy and
metering model validated; leading a customer design session unaccompanied; offshore pod
onboarded to your standards.

Similar jobs