Why This Role Stands Out
You will architect and operate cutting-edge High Performance Computing clusters, driving innovation in GPU-accelerated workloads. This permanent role is perfect for experienced Kubernetes Engineers with a passion for complex systems and a desire to develop advanced automation solutions. Apply now to join a forward-thinking team and shape the future of HPC in Dallas.
Quick Overview
Job Description
NOTE-THIS IS A PERM ROLE WITH MY DIRECT CLIENT IN DALLAS AND ONSITE FROM DAY 1. NO 3RD PARTIES, NO VISA CANDIDATES.
Role: Sr. Kubernetes Engineer (HPC, GPU, NVIDIA) Job Location: Dallas, TX Duration: Permanent Hire/Direct Hire Work Model: Onsite Interview: MS Teams Video Education: Bachelor's Degree
Requirements:
- Strong experience using Kubernetes in production environments.
- Experience working with NVIDIA GPUs and Kubernetes.
- Experience with NVIDIA GPU Operator, device plugins, NVML, MIG, and DCGM.
- Proficiency in Go or Python for developing Kubernetes operators and controllers.
- Strong understanding of Kubernetes internals, including CRDs, RBAC, custom controllers, and scheduler extensions.
- Experience supporting GPU-intensive workloads such as large language models (LLMs), training pipelines, and scientific computing.
- Hands-on experience with Helm, Kustomize, and GitOps workflows.
- Familiarity with CNI plugins, especially NVIDIA CNI and Multus.
- Experience monitoring GPU metrics and cluster health using Prometheus and DCGM Exporter.
Responsibilities:
- Architecting and operating Kubernetes clusters optimized for GPU workloads, leveraging NVIDIA GPU Operator, Network Operator and DCGM.
- Developing, deploying and maintaining custom Kubernetes operators and controllers to automate infrastructure services.
- Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features into the scheduling layer.
- Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance.
- Driving observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter and OpenTelemetry.
- Implementing secure multi-user and multi-namespace GPU isolation, with RBAC and policy enforcement, such as OPA or Gatekeeper.
- Maintaining CI/CD pipelines for Kubernetes infrastructure using GitOps, ArgoCD and FluxCD.
- Contributing to infrastructure-as-code, using Terraform, Helm, and Kustomize.
Share resume at
Similar jobs
- GO
Systems Architect
NewGovcio LLC
Fort George G Meade, Maryland🇺🇸Hybrid3 minutes agoTechnology - VA
Full Stack Software Development Engineer
NewVantor
Herndon, Virginia🇺🇸Hybrid3 minutes agoAgileJavaScriptTechnology - LE
Sr. Technical Writer
NewLeidos
Montgomery Village, MD🇺🇸$82.5k - $149.2k/yrRemoteYesterdayAgileCDNConfluence+2Technology - SR
Technical Writer III
NewScientific Research Corporation
Eglin Air Force Base, FL🇺🇸Hybrid2 days agoCDNTechnology - CG
Technical Writer
CGI INC.
Potomac, MD🇺🇸$47/hrHybrid1 week agoMicrosoft OfficeTechnology - CG
Technical Writer
CGI INC.
Laurel, MD🇺🇸$47/hrHybrid1 week agoMicrosoft OfficeTechnology