Quick Overview
Job Description
Duration: Full-time/Permanent
Location: Hybrid, 3 days/week onsite in Downtown Dallas, TX
Our client is expanding its engineering team to build and operate Kubernetes infrastructure supporting large-scale compute and AI workloads. We are seeking a hands-on Kubernetes Platform Engineer with deep experience building, developing, and operating Kubernetes platforms and services used by other engineering teams.
This role is highly technical and approximately 60% hands-on coding, with a strong focus on Kubernetes platform development, Linux, bare-metal infrastructure, automation, distributed systems, and GPU/accelerator technologies.
Responsibilities:
- Build, develop, maintain, and operate Kubernetes platforms and platform services for internal engineering teams.
- Write production-quality code and automation supporting Kubernetes platform capabilities and operations.
- Troubleshoot complex issues across Kubernetes, Linux, networking, and underlying infrastructure.
- Deploy and operate Kubernetes infrastructure on bare-metal servers and Linux systems.
- Work with GPU operators and accelerator infrastructure supporting large-scale compute and AI workloads.
- Develop tools, services, and automation that improve platform reliability and engineering productivity.
- Partner with application and infrastructure engineering teams to understand platform requirements and deliver reliable Kubernetes services.
- Participate in platform improvements, operational troubleshooting, and infrastructure scaling.
- Help evolve Kubernetes infrastructure as the client's data-center and AI compute environment expands.
Required Skills:
- Strong hands-on experience with Kubernetes platform engineering and administration.
- Experience building or owning Kubernetes services/platforms used by other engineering teams.
- Strong software development and coding skills with recent production coding experience.
- Experience with Linux systems administration and troubleshooting.
- Strong understanding of networking and distributed systems.
- Experience troubleshooting complex Kubernetes and infrastructure issues.
- Experience with automation, scripting, and infrastructure tooling.
- Ability to work independently, take ownership, and collaborate effectively with engineering teams.
- Ability to learn and work with unfamiliar technologies and infrastructure.
Preferred Skills:
- Kubernetes running on bare-metal infrastructure.
- Experience with NVIDIA GPUs, GPU Operators, CUDA, or accelerator infrastructure.
- Experience supporting AI/ML or high-performance compute workloads.
- Experience building managed Kubernetes platforms or internal developer platforms.
- Kubernetes certifications such as Certified Kubernetes Administrator (CKA).
- Experience with Kubernetes operators, controllers, CRDs, or platform services.
- Experience with containerization, Linux networking, and infrastructure automation.
Similar jobs
- ET
Senior DevOps Cloud Engineer
NewEsvee Technologies Inc
Plano, TX🇺🇸Hybrid23 hours agoAzureBashGit+8Technology - NI
Devops Engineer
NewNimbusAITech LLC
Columbus, OH🇺🇸On-site23 hours agoSAMLSSOBash+1Technology - GL
Senior Cloud DevOps Engineer
NewGXO Logistics
Fort Worth, Texas🇺🇸Hybrid52 minutes agoDockerShellAWS+15Technology - VA
Platform Engineer
NewVantor
Herndon, Virginia🇺🇸Hybrid52 minutes agoDockerGCPAWS+12Technology - DT
Senior AI Platform Engineer - Nearshore
NewDiligente Technologies
Toronto, ON🇺🇸Hybrid23 hours agoSSOHelmJWT+3Technology - IF
Senior DevOps Engineer
NewiFusion Inc.
Washington, DC🇺🇸On-site23 hours agoSAFeScrumAgile+2Technology