Sr. Staff Site Reliability/SRE _SanJose, CA(100% onsite)_$65 per hour on C2C_Face to Face interview is must (Only Local Candidate, no relocation)
Why This Role Stands Out
This remote Sr. Staff SRE role offers a fantastic opportunity to leverage your extensive experience in cloud-native infrastructure and AI/ML workloads, with a competitive hourly rate of $70. You'll thrive here if you are a seasoned professional who excels at building and maintaining complex distributed systems and is passionate about driving innovation. Apply today to join a dynamic team and make a significant impact!
Quick Overview
Job Description
Xoriant is an equal opportunity employer. No person shall be excluded from consideration for employment because of race, ethnicity, religion, caste, gender, gender identity, sexual orientation, marital status, national origin, age, disability or veteran status.
[**** NOTE - NO RELOCATION CANDIDATE , ONLY LOCAL TO SanJose, CA , Who is ready to attend face to face interview*********]
TITLE:- Sr. Staff SRE (AI/ML)
LOCATION - SanJose, CA
DURATION 12+ Months (May get extend)
MODE OF INTERVIEW- Face to Face must
RATE - $65 per hour on C2C
JOB DESCRIPTION
- Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
Terraform or OpenTofu proficiency.
Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
Strong automation skills in Python, Bash, or Go.
Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
Proven ability to troubleshoot complex distributed systems, largely self-directed.
Preferred Qualifications GPU infrastructure and AI/ML workloads: Ray, Kubeflow, ML flow, or similar.
NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
Distributed tracing and Open Telemetry instrumentation across services.
Progressive delivery: canary and blue/green rollouts with automated rollback.
Chaos or fault-injection testing, game days, and disaster-recovery drills.
Multi-cloud networking, unified storage abstractions, and disaster recovery.
FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
Establishing an SRE function where one did not previously exist.
Similar jobs
- KE
SAP S/4HANA Public Cloud Infrastructure/DevOps Lead | Remote | Must: S/4HANA Public Cloud, BTP, CBC, Cloud ALM, CI/CD, Git, 3SL, PPL, ECC/HANA, SAP Architecture & Infrastructure Leadership.
NewKeylent
United States🇺🇸Remote10 hours agoGitTechnology - RT
Software Platform Engineer I (Onsite) with Security Clearance
NewRTX
Marlborough, MA🇺🇸Hybrid10 hours agoMATLABScrumAgile+8Technology - BA
Cloud Engineer Specialist - Cloud/DevOps with Security Clearance
Bogart Associates
Reston, VA🇺🇸Hybrid7 weeks agoSAFeDockerAWS+8Technology - LE
Junior DevOps Engineer
NewLeidos
Aldie, VA🇺🇸$69.5k - $125.7k/yrHybridYesterdayDockerELKScrum+8Technology - LE
Junior DevOps Engineer
NewLeidos
Herndon, VA🇺🇸$69.5k - $125.7k/yrHybridYesterdayDockerELKScrum+8Technology - LE
Junior DevOps Engineer
NewLeidos
Manassas, VA🇺🇸$69.5k - $125.7k/yrHybridYesterdayDockerELKScrum+8Technology