Haystack
← Back to Jobs
Remote
Technology
DD

GPU Solutions Architect - Cloud

Digital Dhara LLCUnited States🇺🇸United StatesPosted 12 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Job Title: GPU Solutions Architect - Cloud

Location: Remote, US (Preferred: Santa Clara, CA)

Locations:

  • Santa Clara, CA
  • Remote (CA, WA, OR, TX)

Duration: 1 Year+

# of Roles: 3

 

Important Details

  • Position can be fully remote within the US.
  • Candidates near Santa Clara, CA are preferred.
  • Overtime may be required, including weekends when necessary.
  • Occasional travel is required.
  • Expected contract duration is 1 year.

Required Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, Physics, or related discipline (or equivalent experience).
  • 8+ years of experience in:
    • Production Infrastructure
    • Cloud Engineering
    • Solutions Architecture
    • Site Reliability Engineering (SRE)
    • HPC Environments
    • Similar technical disciplines OR
  • 5+ years of exceptional specialist-level experience supporting large-scale GPU or AI infrastructure.

Technical Expertise: Experience building, operating, and optimizing distributed infrastructure in production environments. Deep hands-on expertise in one or more of the following:

GPU Infrastructure:

  • DCGM
  • BMC / Redfish
  • Firmware Lifecycle Management
  • Driver Lifecycle Management

Networking

  • InfiniBand
  • High-Speed Ethernet
  • NCCL
  • UFM

High-Performance Storage

  • Lustre
  • IBM Storage Scale (GPFS)
  • WEKA
  • VAST Data
  • Similar Enterprise Storage Platforms

Platform Experience

Hands-on experience with:

  • Kubernetes
  • Slurm
  • GPU Scheduling
  • Multi-Tenancy Architecture

Observability

  • Prometheus
  • Grafana
  • OpenTelemetry

Automation & Infrastructure as Code

  • Terraform
  • Ansible
  • Argo CD
  • Similar Automation Frameworks

Operating Systems & Scripting

  • Linux Administration
  • Python
  • Bash
  • Comparable scripting languages

Soft Skills

  • Strong root-cause analysis and troubleshooting skills.
  • Ability to communicate complex technical findings clearly.
  • Experience leading technical initiatives without direct authority.
  • Strong customer-facing communication skills.
  • Ability to manage multiple partner engagements simultaneously.

Preferred Qualifications

Candidates will stand out if they have:

  • Experience operating GPU cloud environments under production workloads.
  • Experience managing large-scale AI platforms or HPC environments.
  • Built or improved 24x7 operational support functions.
  • Experience with:
    • Observability platforms
    • Incident management
    • Problem management
    • On-call systems

Highly Desired

Hands-on experience with:

  • GB200 NVL72
  • GB300 NVL72
  • Spectrum-X
  • UFM
  • Base Command Manager
  • Mission Control
  • GPU Operators
  • Network Operators

Skills

Ansible
Bash
Grafana
Kubernetes
Prometheus
Python
Terraform

Similar jobs