Why This Role Stands Out
This role offers significant opportunities to deepen your expertise in cutting-edge hardware and infrastructure management within a dynamic data center environment. You'll thrive here if you have a strong background in server and GPU break/fix, OS administration, and scripting, and are eager to contribute to a high-performing technical team. Apply now to advance your career in this exciting opportunity with TrustIT LLC.
Quick Overview
Job Description
Data Center Technician / Engineer
Position Title: Data Center Technician / Engineer
Location: Reno, NV (5 Days Onsite)
Duration: 12 Months Contract
Target Experience: 5 to 8 Years
Core Specialty: GPU & Server Hardware Break/Fix, Linux/Windows OS Administration, DCIM Tooling, Scripting (Python/Shell/Ansible)
High-Value Skills: High-Performance Computing (HPC) clusters (Slurm, Bright Cluster Manager), liquid cooling, dense rack layout, networking protocols (TCP/IP, DNS, NFS, SSL)
Key Responsibilities:
Hardware & Compute Farm Management: Maintain a high-performing compute farm of builders, packagers, testers, and core server infrastructure.
Server & GPU Break/Fix: Perform hands-on troubleshooting and replacement for PCBs, GPUs, power supplies, memory, and high-density compute nodes.
Automation & Scripting: Use Shell, Python, or Ansible to automate recurring tasks, run operational scripts, and manage DCIM tooling (e.g., Nautobot).
Cross-Functional Operations: Collaborate with system architects, software developers, and QA engineers to debug hardware/software edge cases and meet availability SLAs.
Process Documentation: Author Standard Operating Procedures (SOPs), collect key operational metrics, and manage system recovery efforts.
Required Qualifications & Skills:
Associate s or Bachelor s degree in a technical major (or equivalent hands-on experience).
5 to 8 years of direct experience in data center environments or large engineering labs.
Strong operating system administration across Linux, Windows, and macOS.
Hands-on scripting proficiency with Python, Shell, or Ansible.
Working knowledge of network protocols: TCP/IP, DNS, NFS, SSL.
Direct experience with DCIM tools (Nautobot or similar inventory/rack management systems).
Preferred / Standout Skills:
Experience managing HPC clusters using Slurm or Bright Cluster Manager (BCM).
Knowledge of dense server infrastructure, including liquid cooling systems.
Network certifications such as CCNA or equivalent.
Similar jobs
- BT
Business Intelligence Analyst
NewBeck Technology Inc
Dallas, TX🇺🇸RemoteYesterdaySQLTableauAzure+5Technology - BS
IT Support Technician (SEP)
NewBabbitting Service Inc
South Elgin, IL🇺🇸On-siteYesterdayTCP/IPActive DirectoryTechnology - 5H
Marketing Optimization Data Analyst
New51905 HROC LLC
United States🇺🇸Hybrid5 hours agoTechnology - BG
System Administrator/Lead Technician - Physical Security Systems
NewBrookfield Global
San Antonio, TX🇺🇸On-siteYesterdayMicrosoft OutlookTechnology - GL
DevOps Engineer - Plano, TX - W2 Position ONLY
NewGeopaq Logic
Plano, TX🇺🇸HybridYesterdayLLMTechnology - IN
Systems Operations Engineer
NewInnova
Charlotte, NC🇺🇸$40 - $45/hrHybridYesterdayAWSTechnology