Haystack
← Back to Jobs
Remote
Other
BI

HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLCUnited States🇺🇸United StatesPosted Oct 7, 2026

Quick Overview

Salary
$170k/yr
Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
Yesterday
ShellAWSMLflowBashCUDAGoogle CloudHelmKubernetesPythonSchedulingTerraform

Job Description

Position : HPC (High-Performance Computing) Consultant Location: Remote Employment: Full-Time | No Contractors Experience: 10+ Years
Role Summary
We are seeking an experienced HPC / Kubernetes / Cloud Infrastructure Engineer with strong hands-on experience supporting large-scale Kubernetes environments, HPC/GPU infrastructure, AI/ML workloads, cloud platforms, infrastructure automation, and SRE operations.
Mandatory Skills
  • Kubernetes large-scale production environments, cluster lifecycle, node management, upgrades, troubleshooting
  • Kubernetes Troubleshooting Scheduler, CNI/networking, nodes, storage, and cluster issues
  • HPC / High-Performance Computing
  • Slurm
  • NVIDIA GPU Infrastructure and GPU scheduling
  • AI/ML Infrastructure and GPU-based workloads
  • AWS especially EKS, EC2, VPC, IAM, S3, and FSx for Lustre
  • Terraform / Infrastructure as Code (IaC)
  • Python automation and coding
  • Bash/Shell scripting
  • Monitoring & SRE PrometheGrafana or equivalent, alerting, SLI/SLO, incident response, RCA
  • Kubernetes Networking Calico or Cilium
  • Helm, RBAC, autoscaling, and rolling upgrades
  • Experience supporting large-scale Kubernetes clusters (500+ nodes) is strongly preferred.
Preferred Skills
  • Multi-cloud: AWS, Google Cloud Platform, CoreWeave
  • Karpenter
  • CUDA
  • AWS ParallelCluster
  • Lustre / WekaFS
  • InfiniBand / RDMA
  • Kubeflow, KServe, Ray, MLflow, or vLLM
  • Distributed AI/ML training and inference platforms
Ideal Candidate: Strong hands-on Kubernetes + Slurm + NVIDIA GPU + AI/ML Infrastructure + Terraform + Python + AWS/EKS/FSx for Lustre + SRE/Monitoring experience.

Similar jobs