Why This Role Stands Out
This hybrid Site Reliability Engineer role offers an exciting opportunity to work with cutting-edge AI/GPU infrastructure, developing valuable automation and diagnostic skills while collaborating with expert teams. If you thrive on solving complex Linux and hardware challenges and want to contribute to the future of AI, this position is an excellent next step in your career. Apply today to join a dynamic environment and make a significant impact!
Quick Overview
Job Description
We are seeking an experienced Site Reliability Engineer (SRE) for an hourly contract engagement supporting cutting-edge AI/GPU infrastructure. The consultant will troubleshoot complex Linux, server hardware, GPU, and infrastructure issues while collaborating with Engineering, Data Center Operations, and Capacity teams.
Key Responsibilities
- Troubleshoot complex Linux, GPU, server hardware, firmware, and infrastructure issues.
- Analyze system/kernel logs and BMC/Redfish telemetry to identify root causes.
- Support hardware provisioning, repair, deployment, and production validation.
- Develop automation, diagnostics, provisioning, and hardware repair tools.
- Test and validate next-generation AI servers and GPU platforms.
- Develop operational documentation, procedures, and best practices.
- Collaborate with Hardware Engineering, Data Center Operations, and Capacity Planning teams.
- Provide on-call and remote operational support as required.
Required Qualifications
- Strong Linux administration and Linux internals knowledge.
- Proven experience troubleshooting server hardware and infrastructure.
- Strong understanding of GPU, hardware, firmware, networking, and systems troubleshooting.
- Excellent root-cause analysis and problem-solving skills.
- Experience with infrastructure provisioning and operations.
- Strong communication and cross-functional collaboration skills.
- Bachelor's degree in Computer Science or related field, or equivalent experience.
Nice to Have
- Experience with large-scale GPU/AI infrastructure.
- Programming experience with Python, Go, or similar languages.
- Experience developing infrastructure or hardware automation tools.
- Experience with BMC/Redfish and data center hardware operations.
Similar jobs
- CH
Site Reliability Engineer ll
NewCohere Health
United States🇺🇸$100k - $110k/yrHybrid12 hours agoMySQLNode.jsAWS+9Technology - BR
Sr Site Reliability Engineer, Platform
NewBlue River Technology
Remote-US🇺🇸$174k - $305k/yrRemote11 hours agoRustAWSMachine Learning+10Technology - RI
Sr. Data and Platform Engineer
NewRivianvw.tech
Irvine🇺🇸Hybrid10 hours agoSQLAWSTableau+6Technology - ES

Technology Operations Engineer
NewEagle Seven
Chicago, Illinois🇺🇸$75k/yrHybrid45 minutes agoNagiosChefGrafana+5Technology - BR
DevSecOps Engineer
NewBryceTech
Suffolk, VA🇺🇸$130k - $150k/yrOn-site14 hours agoEngineering - FR
Platform Operations Engineer
NewFreedomPay
Philadelphia, Pennsylvania🇺🇸Hybrid1 hour agoSQLT-SQLEncryption+11Technology