Haystack
← Back to Jobs
Technology

Looking for Senior/Staff SRE for AI/ML Platform Infrastructure

Xoriant CorporationSan Jose, CA🇺🇸United StatesPosted 10 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Job Title: Senior/Staff SRE for AI/ML Platform Infrastructure

Location: San Jose, CA (Hybrid)

Duration: 6+ Months (Extension Possible)

 

Minimum Qualifications
Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
Terraform or OpenTofu proficiency.
Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
Strong automation skills in Python, Bash, or Go.
Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
Proven ability to troubleshoot complex distributed systems, largely self-directed.


Preferred Qualifications
GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar.
NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
Distributed tracing and OpenTelemetry instrumentation across services.

Skills

Docker
AWS
MLflow
ArgoCD
Azure
Bash
CUDA
DNS
GitHub Actions
GitLab CI
Google Cloud
Grafana
Jenkins
Kubernetes
Prometheus
Python
Terraform

Similar jobs