Haystack
← Back to Jobs
Technology
XO

AI/HPC System Engineer

Xoriant CorporationSan Jose, CA🇺🇸United StatesPosted 24 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
San Jose, CA, United States
Posted
3 days ago
AWSAzureGoogle CloudKubernetesLLM

Job Description

Position Title: AI/HPC System Engineer

Location: San Jose, CA (Onsite)

 

Description

We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development workloads. This role deploys, automates, and maintains GPU clusters across on-premise and cloud environments, delivering reliable, scalable, and cost-efficient compute for engineering and R&D teams.

 

Responsibilities:

• GPU/HPC infrastructure: Build, configure, and operate GPU and HPC clusters across compute, storage, and networking; support capacity planning, performance tuning, and optimization for AI training, inference, and compute-intensive workloads

• Hybrid cloud infrastructure: Deploy and maintain compute environments spanning on-premise and public cloud, and contribute to modernization and scaling initiatives for HPC/AI infrastructure

• Automation and observability: Implement infrastructure-as-code, provisioning automation, monitoring, and alerting, and drive improvements in resource utilization and efficiency

• AI platform support: Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by internal engineering teams

• Operations and collaboration: Troubleshoot and resolve infrastructure issues, document standards and runbooks, and work with relevant stakeholders to support day-to-day IT operations

 

Qualifications:

• Bachelor''s degree in Computer Science, Engineering, or a related technical field

• 3+ years of hands-on experience in IT infrastructure, cloud, platform engineering, or HPC

• Hands-on experience with Linux-based infrastructure and public cloud environments such as AWS, Azure, or Google Cloud Platform

• Experience deploying or operating GPU/HPC environments, including workload scheduling or orchestration platforms such as Kubernetes or Slurm

• Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization

• Solid understanding of compute, storage, networking, and container technologies; experience with AI/ML infrastructure or workloads is a plus

• Strong collaboration and communication skills, with the ability to work across engineering and IT teams

Similar jobs