Why This Role Stands Out
You can drive innovation by building and automating cutting-edge AI and HPC infrastructure, offering significant growth potential within a reputable tech company. This role is perfect for a proactive engineer who thrives on complex system challenges and wants to make a tangible impact on R&D. Apply today to shape the future of advanced computing!
Quick Overview
Job Description
Position Title: AI/HPC System Engineer
Location: San Jose, CA (Onsite)
Description
We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development workloads. This role deploys, automates, and maintains GPU clusters across on-premise and cloud environments, delivering reliable, scalable, and cost-efficient compute for engineering and R&D teams.
Responsibilities:
• GPU/HPC infrastructure: Build, configure, and operate GPU and HPC clusters across compute, storage, and networking; support capacity planning, performance tuning, and optimization for AI training, inference, and compute-intensive workloads
• Hybrid cloud infrastructure: Deploy and maintain compute environments spanning on-premise and public cloud, and contribute to modernization and scaling initiatives for HPC/AI infrastructure
• Automation and observability: Implement infrastructure-as-code, provisioning automation, monitoring, and alerting, and drive improvements in resource utilization and efficiency
• AI platform support: Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by internal engineering teams
• Operations and collaboration: Troubleshoot and resolve infrastructure issues, document standards and runbooks, and work with relevant stakeholders to support day-to-day IT operations
Qualifications:
• Bachelor''s degree in Computer Science, Engineering, or a related technical field
• 3+ years of hands-on experience in IT infrastructure, cloud, platform engineering, or HPC
• Hands-on experience with Linux-based infrastructure and public cloud environments such as AWS, Azure, or Google Cloud Platform
• Experience deploying or operating GPU/HPC environments, including workload scheduling or orchestration platforms such as Kubernetes or Slurm
• Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization
• Solid understanding of compute, storage, networking, and container technologies; experience with AI/ML infrastructure or workloads is a plus
• Strong collaboration and communication skills, with the ability to work across engineering and IT teams
Similar jobs
- SH
Principal AI Engineer
NewSonic Healthcare USA, Inc
Dallas, Texas🇺🇸Hybrid19 minutes agoDockerGCPSQL+11Technology - EV
Regional AI Business Lead
NewEverpure, Inc.
New York, TX🇺🇸$141.5k/yrHybrid11 hours agoMEDDICSales EnablementSales - IN
AI Voice Evaluation Specialist
NewInnodata Inc.
Remote - Alabama ; Remote - Alaska ; Remote - Arkansas ; Remote - Delaware ; Remote - Florida ; Remote - Georgia ; Remote - Indiana ; Remote - Maryland ; Remote - Massachusetts ; Remote - New Hampshire ; Remote - New Jersey ; Remote - New York ; Remote - South Carolina ; Remote - Texas🇺🇸Remote18 hours agoGenerative AI - DE
Applied Researcher – AI Expert
NewDesignworkstalent
Bellevue🇺🇸Remote14 hours agoMachine LearningCUDACapacity Planning+1 - CO
Lead AI Engineer (FM Hosting, LLM Inference)
NewCapital One
Mc Lean, Virginia🇺🇸$197.3k - $225.1k/yrHybrid13 hours agoScalaAWSMachine Learning+9Technology - CO
Senior Lead AI Engineer (FM Hosting, LLM Inference)
NewCapital One
Mc Lean, Virginia🇺🇸$229.9k - $262.4k/yrHybrid14 hours agoScalaAWSMachine Learning+9Technology