Haystack
← Back to Jobs
Technology
LT

Kubernetes Platform Engineer

Laiba Technologies LLCSanta Clara, CA🇺🇸United StatesPosted 26 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Santa Clara, CA, United States
Posted
22 hours ago
Service MeshBashDNSGrafanaHelmKubernetesPrometheusPython

Job Description

Role: Kubernetes Platform Engineer 
Location: Santa Clara, CA(Onsite)

AI Infrastructure
The Candidate will provide senior Kubernetes platform engineering services for AI infrastructure environments supporting model development, distributed training, inference services, and shared platform operations. The role requires strong cluster troubleshooting ability plus pragmatic platform engineering skills in mixed bare-metal and data center environments.

WHAT THIS CANDIDATE WILL BE DOING
·            Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads.
·            Diagnose failures across control plane components, Kubernetes, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
·            Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
·            Improve platform reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
·            Investigate workload issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
·            Partner with Linux, network, validation, and SRE teams to resolve complex cross-layer failures affecting AI services.
·            Create reusable operational runbooks, dashboards, and health checks for day-2 support.
·            Contribute to platform hardening, tenant readiness, and service-level objectives.

WHAT WE NEED TO SEE
·            7+ years in infrastructure engineering, with deep hands-on Kubernetes administration experience.
·            Strong operational understanding of Kubernetes internals and cluster troubleshooting.
·            Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
·            Experience supporting GPU workloads on Kubernetes in lab, validation, or production settings.
·            Strong Linux administration foundation and understanding of data center network dependencies.
·            Ability to debug issues from symptom to root cause across node, pod, network, storage, and control plane layers.
·            Scripting and automation skill in Python, Bash, or Go.

PREFERRED EXPERIENCE
·            Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies.
·            Familiarity with bare-metal Kubernetes and high-performance storage integration.
·            Exposure to regulated or high-change-control production environments.

Similar jobs