Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
Santa Clara, CA, United States
Posted
20 hours ago
Service MeshBashDNSGrafanaHelmKubernetesPrometheusPython
Job Description
Role :- Senior Kubernetes Platform
Location :- Santa Clara, CA(Onsite)
Long-Term Contrcat
Senior Kubernetes Platform
AI Infrastructure
Position Description
ENGAGEMENT SUMMARY
The Candidate will provide senior Kubernetes platform engineering services for AI infrastructure environments supporting model development, distributed training, inference services, and shared platform operations. The role requires strong cluster troubleshooting ability plus pragmatic platform engineering skills in mixed bare-metal and data center environments.
WHAT THIS CANDIDATE WILL BE DOING
- Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads.
- Diagnose failures across control plane components, Kubernetes, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
- Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
- Improve platform reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
- Investigate workload issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
- Partner with Linux, network, validation, and SRE teams to resolve complex cross-layer failures affecting AI services.
- Create reusable operational runbooks, dashboards, and health checks for day-2 support.
- Contribute to platform hardening, tenant readiness, and service-level objectives.
WHAT WE NEED TO SEE
- 7+ years in infrastructure engineering, with deep hands-on Kubernetes administration experience.
- Strong operational understanding of Kubernetes internals and cluster troubleshooting.
- Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
- Experience supporting GPU workloads on Kubernetes in lab, validation, or production settings.
- Strong Linux administration foundation and understanding of data center network dependencies.
- Ability to debug issues from symptom to root cause across node, pod, network, storage, and control plane layers.
- Scripting and automation skill in Python, Bash, or Go.
PREFERRED EXPERIENCE
- Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies.
- Familiarity with bare-metal Kubernetes and high-performance storage integration.
- Exposure to regulated or high-change-control production environments.
Similar jobs
- GI
Senior SAP Transformation Consultant
Galaxy i Technologies, Inc.
San Jose, CA🇺🇸Hybrid3 weeks agoTechnology - CA
QA Automation Engineer
CaritaTech LLC.
Dallas, TX🇺🇸Hybrid5 weeks agoSQLETLSelenium+5Technology - CS
Axiom Developer
Cynet Systems
Hanover, NJ🇺🇸$55 - $60/hrHybrid6 weeks agoOracleSQLSQL Server+9Technology - AE
Control Room Operator (Part Time)
AECOM
Omaha, NE🇺🇸$25/hrOn-site5 weeks agoTechnology - AE
Control Room Lead/SOP Specialist
AECOM
Omaha, NE🇺🇸$31/hrOn-site5 weeks agoTechnology - AE
Control Room Operator
AECOM
Omaha, NE🇺🇸$25/hrOn-site5 weeks agoTechnology