← Back to Jobs
Remote
Other
CD
Linux Administrator- AI,HPC Infrastructure
Cloud Destinations LLCUnited States🇺🇸United StatesPosted 14 Aug 2026
Quick Overview
Work Type
Remote
Level
Mid Senior
Job Description
Job Title: Sr. Linux Administrator, AI and HPC Infrastructure
Location: Remote – United States
Work Setting: Fully remote within the United States
Engagement: Contract – 12 Months
Role Overview
Seeking a Sr. Linux Administrator, AI and HPC Infrastructure to provide senior Linux administration services across AI and HPC environments. The environment spans GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.
Key Responsibilities
- Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
- Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
- Diagnose failures spanning BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
- Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
- Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
- Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
- Automate repeatable administration and remediation tasks with Bash and Python.
- Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.
Required Qualifications
- 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
- Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
- Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
- Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
- Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
- Strong shell scripting and Python-based automation capability.
- Working knowledge of storage and network dependencies affecting Linux host health.
- Ability to operate independently in ambiguous, high-severity production situations.
Preferred Qualifications
- Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
- Familiarity with GPU telemetry and health monitoring tooling (e.g., DCGM) and high-performance fabric technologies, including telemetry-driven health analysis.
- Experience supporting validation labs or pre-production cluster certification.
Skills
Shell
Ansible
Bash
DNS
Kubernetes
Python
Scheduling
Similar jobs
Insights & Analytics SME
Redrouthu's Inc · United States
1 minute agoAgentic AI & MCP Architect
Apex Systems · United States
1 minute agoOracle Cloud SCM Lead Consultant
Sierra-Cedar, Inc. · United States
1 minute agoCloud Migartion Consultant
Irvine Technology Corporation (ITC) · United States
1 minute agoOracle EBS SCM Architect (O2C) for 6+ months contract; 100% Remote
Vigilant Technologies · United States
1 minute agoSimcorp GAIN -- CONTRACT - REMOTE
Nutech Information Systems · United States
1 minute ago