Haystack
← Back to Jobs
Technology
XO

Sr. Staff Site Reliability/SRE _SanJose, CA(100% onsite)_$65 per hour on C2C_Face to Face interview is must (Only Local Candidate, no relocation)

Xoriant CorporationSan Jose, CA🇺🇸United StatesPosted 25 Aug 2026

Why This Role Stands Out

This remote Sr. Staff SRE role offers a fantastic opportunity to leverage your extensive experience in cloud-native infrastructure and AI/ML workloads, with a competitive hourly rate of $70. You'll thrive here if you are a seasoned professional who excels at building and maintaining complex distributed systems and is passionate about driving innovation. Apply today to join a dynamic team and make a significant impact!

Quick Overview

Salary
$70/hr
Seniority
Mid Senior
Work mode
On Site
Location
San Jose, CA, United States
Posted
1 week ago
DockerAWSArgoCDAzureBashCUDADNSGitHub ActionsGitLab CIGoogle CloudGrafanaJenkinsKubernetesPrometheusPythonTerraform

Job Description

Xoriant is an equal opportunity employer. No person shall be excluded from consideration for employment because of race, ethnicity, religion, caste, gender, gender identity, sexual orientation, marital status, national origin, age, disability or veteran status.

[**** NOTE - NO RELOCATION CANDIDATE , ONLY LOCAL TO  SanJose, CA , Who is ready to attend face to face interview*********]

 

TITLE:- Sr. Staff SRE (AI/ML)

LOCATION - SanJose, CA

DURATION 12+ Months (May get extend)

MODE OF INTERVIEW- Face to Face must

RATE - $65 per hour on C2C

JOB DESCRIPTION

  • Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
    Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
    Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
    Terraform or OpenTofu proficiency.
    Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
    Strong automation skills in Python, Bash, or Go.
    Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
    CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
    Proven ability to troubleshoot complex distributed systems, largely self-directed.
    Preferred Qualifications GPU infrastructure and AI/ML workloads: Ray, Kubeflow, ML flow, or similar.
    NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
    Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
    Distributed tracing and Open Telemetry instrumentation across services.
    Progressive delivery: canary and blue/green rollouts with automated rollback.
    Chaos or fault-injection testing, game days, and disaster-recovery drills.
    Multi-cloud networking, unified storage abstractions, and disaster recovery.
    FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
    Establishing an SRE function where one did not previously exist.

 

Similar jobs