Quick Overview
Job Description
Join a high-growth infrastructure team operating Kubernetes platforms across multiple cloud providers at massive scale. You'll build the systems that power thousands of GPUs, where your code and configurations directly protect thousands of GPU-hours from costly failures. EPAM is where tech talent thrives-building groundbreaking solutions, advancing your skills through world-class learning platforms, and working alongside a global community of problem-solvers to make the future real.
Req# Responsibilities Operate and scale Kubernetes platforms (EKS, GKE, and other distributions) including cluster lifecycle management, node pool optimization, and networking policies during periods of rapid growth Provision and manage HPC infrastructure through CI/CD pipelines spanning AWS, CoreWeave, Google Cloud Platform, OCI, and additional cloud providers Design and maintain job scheduling systems that efficiently allocate GPU compute resources across training and inference workloads Define SLIs/SLOs, build robust monitoring and alerting systems, and actively participate in incident response and post-incident reviews Develop production-quality tooling and automation to support multi-cloud infrastructure operations at scale Collaborate daily with Networking, Storage, Security, and AI/ML platform teams to ensure seamless cross-functional infrastructure delivery Requirements 10+ years of experience in infrastructure engineering, cloud platforms, or high-performance computing environments Expert-level Kubernetes experience at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets Advanced Python skills with a track record of building production-grade tools, not just scripts; experience with Go, Rust, or C++ is a strong plus Daily proficiency in Terraform for writing and reviewing infrastructure as code Working knowledge of core AWS services including EC2, S3, EFS, and FSx for Lustre Strong site reliability engineering background with experience building monitoring, alerting, and incident response practices CKA, CKS certificates are highly preferred
Similar jobs
- CI
DevOps Software Engineer - Zero Trust ICAM
NewCACI International, Inc.
MD🇺🇸$94.4k - $198.2k/yrHybrid2 days agoDockerMicroservicesEncryption+9Technology - OR
Principal Data Systems Software Engineer - SRE - Top Secret Clearance Required
NewOracle Corporation
Herndon, VA🇺🇸$114.6k - $234.6k/yrHybrid13 hours agoDockerOracleTeamCity+9Technology - ES
Lead Platform Engineer/Architect - HPC, Kubernetes
NewEPAM Systems
New York, NY🇺🇸Hybrid13 hours agoRustAWSC+++4Technology - BA
DevSecOps Engineer
BOOZ, ALLEN & HAMILTON, INC.
Fayetteville, NC🇺🇸$77.5k - $176k/yrOn-site2 weeks agoSAFeScrumAgileEngineering - EW
Director, Site Reliability Engineering - Paze
NewEarly Warning Services, LLC
Scottsdale, AZ🇺🇸$173k - $230k/yrHybrid13 hours agoPhoenixTechnology - SA
Lead DevOps Engineer
NewSAIC
MD🇺🇸$200.0k - $240k/yrRemote13 hours agoMongoDBAWSLogstash+7Technology - WO
Sr Software Development Engineer, SRE (US Federal) with Security Clearance
NewWorkday
Reston, VA🇺🇸$163.8k/yrHybridYesterdayDockerAWSSplunk+8Technology - AG
Azure Cloud Platform Engineer with Security Clearance
NewAgensys Corporation
Houston, TX🇺🇸HybridYesterdayAzureGitHub ActionsTerraformTechnology - PS
Mid-Level DevOps Software Engineer, SE2 with Security Clearance
NewPower3 Solutions
Hanover, MD🇺🇸$206k - $239k/yrHybridYesterdayMongoDBLogstashAgile+9Technology - XT
Senior Systems Engineer/Senior DevOps Engineer with Security Clearance
NewX Technologies, Inc
San Antonio, TX🇺🇸HybridYesterdayDockerAnsibleBash+5Technology - ON
Site Reliability Engineer II
NewAuto ApplyOnapsis
Dallas🇺🇸Hybrid19 hours agoOracleAWSTDD+10Technology - NE
Site Reliability Engineer
NewAuto ApplyNebius
Remote - United States🇺🇸$130k - $180k/yrRemote11 hours agoBashPythonTechnology