Haystack
← Back to Jobs
Remote
Technology
BI

Senior Kubernetes Platform Engineer Multi-Cloud & HPC @ US Remote - Full Time

BURGEON IT SERVICES LLCUnited States🇺🇸United StatesPosted Sep 23, 2026

Why This Role Stands Out

Advance your expertise in multi-cloud Kubernetes and High-Performance Computing with this remote role, offering significant impact and complex challenges. You'll thrive if you have extensive experience operating large-scale Kubernetes clusters and a passion for infrastructure-as-code. Apply today to shape critical infrastructure for cutting-edge workloads.

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
6 days ago
AWSGoogle CloudKubernetesPythonTerraform

Job Description

Senior Kubernetes Platform Engineer Multi-Cloud & HPC
Location: USA _ Remote
Duration: Full Time

This role is Kubernetes-heavy. You'll operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that novel failure modes are routine.
Responsibilities

Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You're responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth.

Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be expanded in the near future.

Manage job scheduling to allocate GPU compute across training and inference workloads.

Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews.

Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.

Requirements

4+ years in infrastructure engineering, cloud platforms, or HPC.

Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.

Terraform proficiency. You'll write and review infrastructure-as-code daily.

Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).

Python for tooling and automation.

Similar jobs