Quick Overview
Job Description
Role :- Senior Linux Administration
Location :- Santa Clara, CA(Onsite)
Duration: Long Term Contract
AI and HPC Infrastructure
Description
Engagement Summary
The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.
WHAT THIS CANDIDATE WILL BE DOING
· Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
· Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
· Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
· Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
· Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
· Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
· Automate repeatable administration and remediation tasks with Bash and Python.
· Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.
WHAT WE NEED TO SEE
· 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
· Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
· Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
· Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
· Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
· Strong shell scripting and Python-based automation capability.
· Working knowledge of storage and network dependencies affecting Linux host health.
· Ability to operate independently in ambiguous, high-severity production situations.
PREFERRED EXPERIENCE
· Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
· Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.
Experience supporting validation labs or pre-production cluster certification
Similar jobs
- KT
Urgent need- React Developers in Seattle, WA
NewKomplete Teleservices
Seattle, WA🇺🇸Hybrid11 hours agoGitHub ActionsJavaScriptJenkins+4 - FB
Director Application Development & Support
First-Citizens Bank & Trust Company
NC🇺🇸Remote1 week agoMicroservicesAgileCompliance+10 - AC
Epic Ambulatory / Security / Kaleidoscope Analyst ** 100% Remote **
Amerit Consulting
United States🇺🇸Remote3 weeks agoRecruiting - DC
Oracle fusion techno-Functional Consultant
Data Capital Inc
Durham, NC🇺🇸Hybrid5 weeks agoOracleSOAPSQL+5 - HI
Integrated Training Systems Network Communications Engineer (Network Communications 2)
NewHII
Virginia Beach, VA🇺🇸$66.3k - $80k/yrHybrid2 days agoMachine LearningTCP/IPAuditing+9 - BS
Maven Exploitation Specialist/Imagery Scientist (SAR Focused) EX with Security Clearance
BTS Software Solutions
Springfield, VA🇺🇸$210k - $230k/yrHybrid5 weeks ago401kMATLABETL+3