Haystack
← Back to Jobs
Technology

GPU Infrastructure Site Reliability Engineer(On-site, L2- Face to Face)

Balin Technologies LLCSunnyvale, CA🇺🇸United StatesPosted 5 Aug 2026

Why This Role Stands Out

You will play a vital role in maintaining cutting-edge AI and GPU infrastructure, offering significant opportunities for technical growth and impact within a reputable company. This on-site position is perfect for experienced SREs who thrive on complex problem-solving and ensuring the reliability of high-performance systems. Apply now to contribute to critical technology advancements.

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Job Title: Site Reliability Engineer (SRE)

Location: Sunnyvale, CA (On-site)
Duration: Long-Term Contract

Job Description

We are seeking a highly motivated Site Reliability Engineer (SRE) to support mission-critical AI and GPU infrastructure in a high-performance production environment. The ideal candidate will have experience supporting GPU platforms, embedded infrastructure, compute, networking, and storage systems while ensuring high availability, reliability, and operational excellence.

As an SRE, you will monitor production environments, troubleshoot infrastructure issues, respond to incidents, and collaborate with platform, hardware, and engineering teams to maintain scalable and reliable infrastructure.

Key Responsibilities

  • Monitor and maintain production infrastructure supporting GPU and embedded platforms.
  • Investigate and resolve infrastructure incidents, system alerts, and performance issues.
  • Support GPU servers, compute infrastructure, storage systems, and networking components.
  • Perform troubleshooting across Linux systems, hardware, networking, storage, and platform services.
  • Work closely with Platform Engineering, Infrastructure, Network, and Hardware teams to ensure platform reliability.
  • Support deployment, provisioning, configuration, and maintenance of infrastructure components.
  • Monitor infrastructure health and proactively identify reliability and performance issues.
  • Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
  • Participate in infrastructure upgrades, maintenance activities, and production rollouts.
  • Create and maintain operational documentation, runbooks, and standard operating procedures.
  • Participate in on-call rotation and provide production support for critical infrastructure.

Required Skills

  • Strong experience in Site Reliability Engineering (SRE) or Infrastructure Engineering.
  • Hands-on experience supporting GPU-based infrastructure.
  • Experience with Embedded Platform Engineering environments.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of compute infrastructure.
  • Experience with enterprise infrastructure and production operations.
  • Knowledge of networking concepts with experience in OVS (Open vSwitch) and OCS.
  • Experience with enterprise storage solutions such as Lightbits and Pure Storage.
  • Understanding of GKN Compute environments or similar compute platforms.
  • Experience in infrastructure monitoring, incident management, and alert handling.
  • Strong troubleshooting and root cause analysis skills.
  • Excellent communication and collaboration skills.

Preferred Skills

  • Experience with Kubernetes or container platforms.
  • Knowledge of cloud platforms (AWS, Azure, or Google Cloud Platform).
  • Familiarity with automation using Bash, Python, or Ansible.
  • Experience with monitoring tools such as Prometheus, Grafana, Datadog, or Splunk.
  • Knowledge of CI/CD and Infrastructure as Code (Terraform, Ansible).

Mandatory Skills

  • GPU Infrastructure
  • Embedded Platform Engineering
  • Linux Administration
  • Infrastructure Operations
  • Compute (GKN or similar)
  • OVS (Open vSwitch)
  • OCS Networking
  • Lightbits Storage
  • Pure Storage
  • Incident Management
  • Alert Monitoring
  • Root Cause Analysis
  • Production Support
  • Troubleshooting

Skills

AWS
Splunk
Ansible
Azure
Bash
Datadog
Google Cloud
Grafana
Kubernetes
Prometheus
Python
Terraform

Similar jobs