Haystack
← Back to Jobs
Technology

Senior/Staff SRE for AI/ML Platform Infrastructure

Xoriant CorporationSan Jose, CA🇺🇸United StatesPosted 11 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Minimum Qualifications
 Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
 Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
 Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
 Terraform or OpenTofu proficiency.
 Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
 Strong automation skills in Python, Bash, or Go.
 Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
 CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
 Proven ability to troubleshoot complex distributed systems, largely self-directed.


Preferred Qualifications
 GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar.
 NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
 Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
 Distributed tracing and OpenTelemetry instrumentation across services.
 Progressive delivery: canary and blue/green rollouts with automated rollback.
 Chaos or fault-injection testing, game days, and disaster-recovery drills.
 Multi-cloud networking, unified storage abstractions, and disaster recovery.
 FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
 Establishing an SRE function where one did not previously exist.

 

Thanks & Regards,

Narendra Kunware

Skills

Docker
AWS
MLflow
ArgoCD
Azure
Bash
CUDA
DNS
GitHub Actions
GitLab CI
Google Cloud
Grafana
Jenkins
Kubernetes
Prometheus
Python
Terraform

Similar jobs