Looking for Senior/Staff SRE for AI/ML Platform Infrastructure
Quick Overview
Job Description
Job Title: Senior/Staff SRE for AI/ML Platform Infrastructure
Location: San Jose, CA (Hybrid)
Duration: 6+ Months (Extension Possible)
Minimum Qualifications
• Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
• Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
• Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
• Terraform or OpenTofu proficiency.
• Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
• Strong automation skills in Python, Bash, or Go.
• Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
• CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
• Proven ability to troubleshoot complex distributed systems, largely self-directed.
Preferred Qualifications
• GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar.
• NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
• Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
• Distributed tracing and OpenTelemetry instrumentation across services.
Skills
Similar jobs
Network Engineer
SmallArc, Inc · Edison, United States
26 minutes agoNetwork Cloud Terraform DevOps Engineer Lead
InfiCare · Palo Alto, United States
26 minutes agoSenior Devops Engineer/Lead
Infinite Computer Solutions (ICS) · Dallas, United States
32 minutes ago$80k/yrPrincipal Engineer Software - DevOps with Security Clearance
Northrop Grumman · Melbourne, United States
2 hours ago$98.4k - $147.6k/yrRelease Engineer (CI/CD, Infrastructure & Automation)
Openmind Technologies · United States
3 hours agoDevOps Engineer [AQ-12187]
Aquent Talent · Phoenix, United States
5 hours ago