Haystack
← Back to Jobs
Other
PT

Senior Linux Administration

Prudent Technologies and ConsultingSanta Clara, CA🇺🇸United StatesPosted 26 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Santa Clara, CA, United States
Posted
Yesterday
ShellAnsibleBashDNSKubernetesPythonScheduling

Job Description

Role :- Senior Linux Administration

Location :- Santa Clara, CA(Onsite)

Duration: Long Term Contract

AI and HPC Infrastructure

Description

Engagement Summary

The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.

WHAT THIS CANDIDATE WILL BE DOING

·                Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.

·                Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.

·                Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.

·                Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.

·                Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.

·                Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.

·                Automate repeatable administration and remediation tasks with Bash and Python.

·                Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.

WHAT WE NEED TO SEE

·                7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.

·                Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.

·                Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.

·                Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.

·                Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.

·                Strong shell scripting and Python-based automation capability.

·                Working knowledge of storage and network dependencies affecting Linux host health.

·                Ability to operate independently in ambiguous, high-severity production situations.

PREFERRED EXPERIENCE

·                Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.

·                Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.

Experience supporting validation labs or pre-production cluster certification

Similar jobs