Haystack
← Back to Jobs
Remote
Other
LO

HPC-High Performance Computing Consultant

Logicplanet, Inc.United States🇺🇸United StatesPosted Oct 8, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
19 hours ago
OracleShellAWSMLflowSplunkAnsibleAzureBashCUDACapacity PlanningDatadogGitHub ActionsGitLab CIGoogle CloudGrafanaHelmJenkinsKubernetesPrometheusPythonRESTRoot Cause AnalysisSchedulingTerraform

Job Description

Role: HPC (High-Performance Computing) Consultant (5+ Openings)

Location: Remote

Duration: Fulltime 

Must Have: Kubernetes + Slurm + NVIDIA GPU + AI/ML Infrastructure + Terraform + Python + AWS (EKS/FSx Lustre) + Monitoring/SRE

 

We are looking for engineers with expertise across Kubernetes, cloud infrastructure, HPC platforms, GPU computing, Terraform, and automation. Depending on experience, candidates may be considered for Platform Engineer, Kubernetes Engineer, HPC Engineer, Cloud Infrastructure Engineer, DevOps Engineer, AI Infrastructure Engineer, or Site Reliability Engineer (SRE) roles.

 

Key Responsibilities

  • Design, deploy, operate, and support large-scale Kubernetes platforms across AWS, Google Cloud Platform, CoreWeave, OCI, and other cloud environments.
  • Manage Kubernetes cluster lifecycle activities including provisioning, scaling, node pool management, upgrades, troubleshooting, and performance optimization.
  • Support AI/ML and HPC workloads, including GPU-enabled compute infrastructure for training and inference environments.
  • Provision and automate cloud and infrastructure resources using Terraform and CI/CD pipelines.
  • Troubleshoot Kubernetes scheduling, networking, storage, and platform reliability issues.
  • Implement monitoring, observability, alerting, SLIs/SLOs, and incident response processes.
  • Collaborate with Networking, Security, Storage, AI/ML, Data Engineering, and Application teams.
  • Develop automation and operational tooling using Python and cloud-native technologies.
  • Participate in production support, root cause analysis, capacity planning, and platform optimization initiatives.

 

Required Skills

Kubernetes & Container Platforms

  • Kubernetes (EKS, GKE, AKS, OpenShift, CoreWeave CKS)
  • Cluster Lifecycle Management
  • Node Pool Management
  • Scheduler Troubleshooting
  • CNI Troubleshooting
  • Networking Policies
  • RBAC
  • Helm
  • Autoscaling
  • Rolling Upgrades

 

Cloud Infrastructure

  • AWS (EC2, S3, IAM, VPC, EKS, EFS, FSx for Lustre)
  • Google Cloud Platform (Google Cloud Platform)
  • OCI (Oracle Cloud Infrastructure)
  • Azure (Preferred)
  • Multi-Cloud Infrastructure

 

Infrastructure as Code & Automation

  • Terraform
  • Infrastructure as Code (IaC)
  • CI/CD Pipelines
  • GitHub Actions / Jenkins / GitLab CI
  • Ansible (Preferred)

 

Programming & Scripting

  • Python
  • Bash/Shell Scripting
  • Automation Development
  • REST API Integration

 

HPC & GPU Infrastructure (Preferred)

  • High-Performance Computing (HPC)
  • Slurm
  • NVIDIA GPU Platforms
  • GPU Scheduling
  • CUDA
  • Distributed Computing
  • AI/ML Infrastructure

 

Monitoring & Reliability

  • Prometheus
  • Grafana
  • Datadog
  • Splunk
  • Cloud Monitoring
  • SLI/SLO Management
  • Incident Response
  • Root Cause Analysis (RCA)

 

Preferred Experience

Experience with any of the following is highly desirable:

  • AI/ML Infrastructure Platforms
  • Kubeflow
  • KServe
  • Ray
  • MLflow
  • vLLM
  • Vector Databases
  • Distributed Training Platforms
  • CoreWeave
  • AWS ParallelCluster
  • FSx for Lustre
  • Lustre
  • WekaFS
  • InfiniBand
  • RDMA
  • AI Training & Inference Workloads

Similar jobs