Haystack
← Back to Jobs
Technology
TB

PERM Role-Direct Client-HPC Systems Engineer (High Performance Computing)

The Brixton GroupDallas, TX🇺🇸United StatesPosted 9 Sept 2026

Why This Role Stands Out

You will architect and operate cutting-edge High Performance Computing clusters, driving innovation in GPU-accelerated workloads. This permanent role is perfect for experienced Kubernetes Engineers with a passion for complex systems and a desire to develop advanced automation solutions. Apply now to join a forward-thinking team and shape the future of HPC in Dallas.

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Dallas, TX, United States
Posted
11 hours ago
ArgoCDGrafanaHelmKubernetesPrometheusPythonTerraform

Job Description

NOTE-THIS IS A PERM ROLE WITH MY DIRECT CLIENT IN DALLAS AND ONSITE FROM DAY 1. NO 3RD PARTIES, NO VISA CANDIDATES.

Role: Sr. Kubernetes Engineer (HPC, GPU, NVIDIA) Job Location: Dallas, TX Duration: Permanent Hire/Direct Hire Work Model: Onsite Interview: MS Teams Video Education: Bachelor's Degree

Requirements:

  • Strong experience using Kubernetes in production environments.
  • Experience working with NVIDIA GPUs and Kubernetes.
  • Experience with NVIDIA GPU Operator, device plugins, NVML, MIG, and DCGM.
  • Proficiency in Go or Python for developing Kubernetes operators and controllers.
  • Strong understanding of Kubernetes internals, including CRDs, RBAC, custom controllers, and scheduler extensions.
  • Experience supporting GPU-intensive workloads such as large language models (LLMs), training pipelines, and scientific computing.
  • Hands-on experience with Helm, Kustomize, and GitOps workflows.
  • Familiarity with CNI plugins, especially NVIDIA CNI and Multus.
  • Experience monitoring GPU metrics and cluster health using Prometheus and DCGM Exporter.

Responsibilities:

  • Architecting and operating Kubernetes clusters optimized for GPU workloads, leveraging NVIDIA GPU Operator, Network Operator and DCGM.
  • Developing, deploying and maintaining custom Kubernetes operators and controllers to automate infrastructure services.
  • Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features into the scheduling layer.
  • Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance.
  • Driving observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter and OpenTelemetry.
  • Implementing secure multi-user and multi-namespace GPU isolation, with RBAC and policy enforcement, such as OPA or Gatekeeper.
  • Maintaining CI/CD pipelines for Kubernetes infrastructure using GitOps, ArgoCD and FluxCD.
  • Contributing to infrastructure-as-code, using Terraform, Helm, and Kustomize.

Share resume at

Similar jobs