Haystack
← Back to Jobs
Remote
Other
CD

Linux Administrator- AI,HPC Infrastructure

Cloud Destinations LLCUnited States🇺🇸United StatesPosted 14 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Job Title: Sr. Linux Administrator, AI and HPC Infrastructure
Location: Remote – United States
Work Setting: Fully remote within the United States
Engagement: Contract – 12 Months
Role Overview
Seeking a Sr. Linux Administrator, AI and HPC Infrastructure to provide senior Linux administration services across AI and HPC environments. The environment spans GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.
Key Responsibilities
  • Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
  • Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
  • Diagnose failures spanning BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
  • Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
  • Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
  • Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
  • Automate repeatable administration and remediation tasks with Bash and Python.
  • Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues. 
Required Qualifications
  • 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
  • Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
  • Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
  • Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
  • Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
  • Strong shell scripting and Python-based automation capability.
  • Working knowledge of storage and network dependencies affecting Linux host health.
  • Ability to operate independently in ambiguous, high-severity production situations. 
Preferred Qualifications
  • Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
  • Familiarity with GPU telemetry and health monitoring tooling (e.g., DCGM) and high-performance fabric technologies, including telemetry-driven health analysis.
  • Experience supporting validation labs or pre-production cluster certification.

Skills

Shell
Ansible
Bash
DNS
Kubernetes
Python
Scheduling

Similar jobs