GPU Infrastructure Site Reliability Engineer(On-site, L2- Face to Face)
Why This Role Stands Out
You will play a vital role in maintaining cutting-edge AI and GPU infrastructure, offering significant opportunities for technical growth and impact within a reputable company. This on-site position is perfect for experienced SREs who thrive on complex problem-solving and ensuring the reliability of high-performance systems. Apply now to contribute to critical technology advancements.
Quick Overview
Job Description
Job Title: Site Reliability Engineer (SRE)
Location: Sunnyvale, CA (On-site)
Duration: Long-Term Contract
Job Description
We are seeking a highly motivated Site Reliability Engineer (SRE) to support mission-critical AI and GPU infrastructure in a high-performance production environment. The ideal candidate will have experience supporting GPU platforms, embedded infrastructure, compute, networking, and storage systems while ensuring high availability, reliability, and operational excellence.
As an SRE, you will monitor production environments, troubleshoot infrastructure issues, respond to incidents, and collaborate with platform, hardware, and engineering teams to maintain scalable and reliable infrastructure.
Key Responsibilities
- Monitor and maintain production infrastructure supporting GPU and embedded platforms.
- Investigate and resolve infrastructure incidents, system alerts, and performance issues.
- Support GPU servers, compute infrastructure, storage systems, and networking components.
- Perform troubleshooting across Linux systems, hardware, networking, storage, and platform services.
- Work closely with Platform Engineering, Infrastructure, Network, and Hardware teams to ensure platform reliability.
- Support deployment, provisioning, configuration, and maintenance of infrastructure components.
- Monitor infrastructure health and proactively identify reliability and performance issues.
- Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
- Participate in infrastructure upgrades, maintenance activities, and production rollouts.
- Create and maintain operational documentation, runbooks, and standard operating procedures.
- Participate in on-call rotation and provide production support for critical infrastructure.
Required Skills
- Strong experience in Site Reliability Engineering (SRE) or Infrastructure Engineering.
- Hands-on experience supporting GPU-based infrastructure.
- Experience with Embedded Platform Engineering environments.
- Strong Linux administration and troubleshooting skills.
- Good understanding of compute infrastructure.
- Experience with enterprise infrastructure and production operations.
- Knowledge of networking concepts with experience in OVS (Open vSwitch) and OCS.
- Experience with enterprise storage solutions such as Lightbits and Pure Storage.
- Understanding of GKN Compute environments or similar compute platforms.
- Experience in infrastructure monitoring, incident management, and alert handling.
- Strong troubleshooting and root cause analysis skills.
- Excellent communication and collaboration skills.
Preferred Skills
- Experience with Kubernetes or container platforms.
- Knowledge of cloud platforms (AWS, Azure, or Google Cloud Platform).
- Familiarity with automation using Bash, Python, or Ansible.
- Experience with monitoring tools such as Prometheus, Grafana, Datadog, or Splunk.
- Knowledge of CI/CD and Infrastructure as Code (Terraform, Ansible).
Mandatory Skills
- GPU Infrastructure
- Embedded Platform Engineering
- Linux Administration
- Infrastructure Operations
- Compute (GKN or similar)
- OVS (Open vSwitch)
- OCS Networking
- Lightbits Storage
- Pure Storage
- Incident Management
- Alert Monitoring
- Root Cause Analysis
- Production Support
- Troubleshooting
Skills
Similar jobs
Network Engineer
TEKsystems c/o Allegis Group · Indian Head, United States
19 minutes ago$45 - $50/hrPrincipal Site Reliability Engineer
Navy Federal Credit Union · Pensacola, United States
1 hour agoLead Site Reliability Engineer(only W2)
Analytics Solutions · Orlando, United States
1 hour agoEmbedded Platform Engineer
Balin Technologies LLC · San Jose, United States
1 hour agoSenior Cloud Platform Engineer
Avanciers LLC · United States
1 hour agoPrincipal Site Reliability Engineer
Navy Federal Credit Union · Vienna, United States
1 hour ago