Haystack
← Back to Jobs
Technology
NE

Site Reliability Engineer (SRE) GPU Infrastructure

Neuveon IncSanta Clara, CA🇺🇸United StatesPosted Sep 25, 2026

Why This Role Stands Out

This hybrid Site Reliability Engineer role offers an exciting opportunity to work with cutting-edge AI/GPU infrastructure, developing valuable automation and diagnostic skills while collaborating with expert teams. If you thrive on solving complex Linux and hardware challenges and want to contribute to the future of AI, this position is an excellent next step in your career. Apply today to join a dynamic environment and make a significant impact!

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Santa Clara, CA, United States
Posted
5 days ago
Python

Job Description

We are seeking an experienced Site Reliability Engineer (SRE) for an hourly contract engagement supporting cutting-edge AI/GPU infrastructure. The consultant will troubleshoot complex Linux, server hardware, GPU, and infrastructure issues while collaborating with Engineering, Data Center Operations, and Capacity teams.

Key Responsibilities

  • Troubleshoot complex Linux, GPU, server hardware, firmware, and infrastructure issues.
  • Analyze system/kernel logs and BMC/Redfish telemetry to identify root causes.
  • Support hardware provisioning, repair, deployment, and production validation.
  • Develop automation, diagnostics, provisioning, and hardware repair tools.
  • Test and validate next-generation AI servers and GPU platforms.
  • Develop operational documentation, procedures, and best practices.
  • Collaborate with Hardware Engineering, Data Center Operations, and Capacity Planning teams.
  • Provide on-call and remote operational support as required.

Required Qualifications

  • Strong Linux administration and Linux internals knowledge.
  • Proven experience troubleshooting server hardware and infrastructure.
  • Strong understanding of GPU, hardware, firmware, networking, and systems troubleshooting.
  • Excellent root-cause analysis and problem-solving skills.
  • Experience with infrastructure provisioning and operations.
  • Strong communication and cross-functional collaboration skills.
  • Bachelor's degree in Computer Science or related field, or equivalent experience.

Nice to Have

  • Experience with large-scale GPU/AI infrastructure.
  • Programming experience with Python, Go, or similar languages.
  • Experience developing infrastructure or hardware automation tools.
  • Experience with BMC/Redfish and data center hardware operations.

Similar jobs