Haystack
← Back to Jobs
Other
ME

HPC (High-Performance Computing) Consultant

MetaRPOUnited States🇺🇸United StatesPosted Sep 24, 2026

Why This Role Stands Out

This remote HPC Consultant role at MetaRPO offers exciting challenges in managing large-scale, multi-cloud Kubernetes platforms, providing significant opportunities for skills development in cutting-edge technologies. You'll thrive here if you possess expertise in Kubernetes, Terraform, and AWS, and are eager to contribute to a dynamic team focused on optimizing critical GPU compute resources.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
20 hours ago
AWSGoogle CloudKubernetesPythonSchedulingTerraform

Job Description

HPC (High-Performance Computing) Consultant

Position Name – HPC (High-Performance Computing) Consultant

Type of hiring – Fulltime

Location – Remote USA

Skill

Experience

Kubernetes

Terraform

AWS

Python

HPC/GPU

CI/CD

Monitoring

Troubleshooting

Production Support

Job Description:

This role is Kubernetes-heavy. You'll operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that novel failure modes are routine.

Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You're responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth.
  • Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be expanded in the near future.
  • Manage job scheduling to allocate GPU compute across training and inference workloads.
  • Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews.
  • Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.

 

Requirements

  • 4+ Years in infrastructure engineering, cloud platforms, or HPC.
  • Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.
  • Terraform proficiency. You'll write and review infrastructure-as-code daily.
  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).
  • Python for tooling and automation.

Similar jobs