Why This Role Stands Out
This remote HPC Consultant role at MetaRPO offers exciting challenges in managing large-scale, multi-cloud Kubernetes platforms, providing significant opportunities for skills development in cutting-edge technologies. You'll thrive here if you possess expertise in Kubernetes, Terraform, and AWS, and are eager to contribute to a dynamic team focused on optimizing critical GPU compute resources.
Quick Overview
Job Description
HPC (High-Performance Computing) Consultant
Position Name – HPC (High-Performance Computing) Consultant
Type of hiring – Fulltime
Location – Remote USA
Skill | Experience |
Kubernetes | |
Terraform | |
AWS | |
Python | |
HPC/GPU | |
CI/CD | |
Monitoring | |
Troubleshooting | |
Production Support |
Job Description:
This role is Kubernetes-heavy. You'll operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that novel failure modes are routine.
Responsibilities
- Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You're responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth.
- Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be expanded in the near future.
- Manage job scheduling to allocate GPU compute across training and inference workloads.
- Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews.
- Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.
Requirements
- 4+ Years in infrastructure engineering, cloud platforms, or HPC.
- Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.
- Terraform proficiency. You'll write and review infrastructure-as-code daily.
- Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).
- Python for tooling and automation.
Similar jobs
- SQ
Test Management Coordinator
NewSonora Quest Laboratories
Glendale, Arizona🇺🇸Hybrid12 minutes agoSQLComplianceJira+4 - CW
Chess Instructor | Fall
NewChess Wizards
Washington, DC🇺🇸$50 - $65/hrHybridYesterday - CW
Chess Instructor | Fall
NewChess Wizards
Barnesville, MD🇺🇸$50 - $65/hrHybridYesterday - CW
Chess Instructor | Fall
NewChess Wizards
Abington, MA🇺🇸$50 - $70/hrHybridYesterday - CW
Chess Instructor | Fall
NewChess Wizards
Los Gatos, CA🇺🇸$65 - $80/hrHybridYesterday - CW
Chess Instructor | Fall
NewChess Wizards
DOWNERS GROVE, IL🇺🇸$65 - $75/hrHybridYesterday