Quick Overview
Job Description
Role: Kubernetes Platform Engineer
Location: Santa Clara, CA(Onsite)
AI Infrastructure
The Candidate will provide senior Kubernetes platform engineering services for AI infrastructure environments supporting model development, distributed training, inference services, and shared platform operations. The role requires strong cluster troubleshooting ability plus pragmatic platform engineering skills in mixed bare-metal and data center environments.
WHAT THIS CANDIDATE WILL BE DOING
· Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads.
· Diagnose failures across control plane components, Kubernetes, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
· Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
· Improve platform reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
· Investigate workload issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
· Partner with Linux, network, validation, and SRE teams to resolve complex cross-layer failures affecting AI services.
· Create reusable operational runbooks, dashboards, and health checks for day-2 support.
· Contribute to platform hardening, tenant readiness, and service-level objectives.
WHAT WE NEED TO SEE
· 7+ years in infrastructure engineering, with deep hands-on Kubernetes administration experience.
· Strong operational understanding of Kubernetes internals and cluster troubleshooting.
· Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
· Experience supporting GPU workloads on Kubernetes in lab, validation, or production settings.
· Strong Linux administration foundation and understanding of data center network dependencies.
· Ability to debug issues from symptom to root cause across node, pod, network, storage, and control plane layers.
· Scripting and automation skill in Python, Bash, or Go.
PREFERRED EXPERIENCE
· Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies.
· Familiarity with bare-metal Kubernetes and high-performance storage integration.
· Exposure to regulated or high-change-control production environments.
Similar jobs
- CA
Senior Infrastructure Automation Engineer / SME remote
Calance
United States🇺🇸Remote4 weeks agoAnsiblePowerShellPython+1Technology - PC
Senior DevOps Engineer (with Azure Cloud Exp)
Pyramid Consulting, Inc.
Dallas, TX🇺🇸$65 - $70/hrOn-site2 weeks agoOracleSQLSQL Server+15Technology - AS
Platform Engineer III
Apex Systems
Cincinnati, OH🇺🇸$60 - $68/hrOn-site6 weeks agoDockerElixirPHP+11Technology - AS
Sr Platform Engineer
Apex Systems
Raleigh, NC🇺🇸On-site6 weeks agoTechnology - AV
Agentic AI/ DevSecOps Engineer
AP Ventures
Columbia, MD🇺🇸Hybrid7 weeks agoAgileEngineering - UT
Sr. AI Platform Engineer #26-21438
NewU.S. Tech Solutions Inc.
St. Louis, MO🇺🇸Hybrid22 hours agoAWSAzureC#+4Technology