Why This Role Stands Out
This fully remote Senior SRE role offers a fantastic opportunity to work with cutting-edge AI and GPU infrastructure, significantly impacting a platform used by over a million developers worldwide. You'll thrive here if you enjoy solving complex reliability challenges and possess strong experience in cloud, Kubernetes, and automation, with the chance to shape the future of AI infrastructure. Apply now to grow your career in a dynamic and rapidly scaling environment.
Quick Overview
Job Description
Remote (USA) | Full-Time | Site Reliability Engineer
Join a rapidly growing B2B AI infrastructure company powering large-scale machine learning and AI workloads for more than one million developers worldwide. As a Site Reliability Engineer, you'll help improve the reliability, scalability, and performance of a cloud platform built on Linux, Kubernetes, distributed systems, GPU infrastructure, observability, and automation technologies. This full-time remote opportunity offers the chance to work on critical infrastructure supporting AI applications on a global scale.
As the company continues to scale its AI infrastructure platform, reliability has become a critical business function. This role sits at the center of that effort, partnering with Infrastructure, Product Engineering, and Support teams to improve uptime, strengthen observability, establish SLOs, reduce operational toil through automation, and lead incident response initiatives. The ideal candidate brings experience supporting large-scale production environments and enjoys solving complex reliability challenges while influencing engineering practices across a rapidly growing organization.
This is an opportunity to gain exposure to cutting-edge AI and GPU infrastructure, take ownership of high-impact initiatives, and help shape the reliability strategy of a platform relied upon by more than one million developers.
Required Skills & Experience
5+ years of experience within major public cloud environment like AWS, Google Cloud Platform
Strong Linux systems administration experience
Strong networking fundamentals and troubleshooting skills
Experience supporting containerized environments (Kubernetes preferred)
Experience with monitoring, alerting, and observability tools
Experience defining and managing SLIs, SLOs, and reliability metrics
Incident response and postmortem experience
Scripting or programming experience. Python, Go, Bash, or similar technologies
Distributed systems and failure scenarios
Desired Skills & Experience
Kubernetes
Prometheus, Grafana, or similar monitoring platforms
Experience supporting GPU infrastructure or AI/ML platforms
Infrastructure as Code experience (Terraform preferred)
CI/CD pipeline experience
What You Will Be Doing
Tech Breakdown
40% Linux & Kubernetes Administration
25% Monitoring, Observability & Incident Response
20% Automation & Reliability Engineering
15% Distributed Systems & Cloud Infrastructure
Daily Responsibilities
80% Hands-On Engineering
5% Management Duties
15% Team Collaboration
The Offer
medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
Annual learning and development budget
Flexible PTO
Career Growth Within a Rapidly Scaling AI Infrastructure Company
Applicants must be currently authorized to work in the US on a full-time basis now and in the future.
Sponsorship is not available for this position
#LI-JG2
Similar jobs
- TS
TTG-345 - Palantir Foundry Data Platform Engineer - $339,950 Tot with Security Clearance
NewTTG Solutions Inc.
Herndon, VA🇺🇸$204k - $247k/yrHybrid17 hours agoETLLESSPythonTechnology - TS
TTG-230 - DevOps Infrastructure Automation Engineer - $265,100 T with Security Clearance
NewTTG Solutions Inc.
Chantilly, VA🇺🇸$158k - $192k/yrHybrid17 hours agoDockerSOAPShell+9Technology - V1
Senior Data Platform Engineer
NewVersion 1
United States🇺🇸Hybrid2 hours agoSQLSQL ServerAWS+5Technology - OR
Senior Site Reliability Engineer
NewOracle Corporation
Nashville, TN🇺🇸$81.1k - $187k/yrHybridYesterdayOracleAnsibleBash+4Technology - OR
Principal Site Reliability Engineer
Oracle Corporation
Nashville, TN🇺🇸$84.9k - $209.5k/yrHybrid3 days agoOracleLoad BalancingAnsible+5Technology - SP
Sr. Site Reliability Engineer, Platform Infrastructure
NewSpaceX
Bastrop, TX🇺🇸Hybrid17 hours agoDockerAnsibleKubernetes+3Technology - CI
DevOps Engineer
CACI International, Inc.
Annapolis, MD🇺🇸$103.8k - $218.1k/yrHybrid4 days agoDockerMongoDBMySQL+12Technology - SY
Mid-Level Cloud Platform Engineer [$274k/yr+] with Security Clearance
NewSYSTOLIC
Annapolis Junction, MD🇺🇸$274k/yrHybrid17 hours agoDockerShellAWS+11Technology - BS
Senior DevOps Software Engineer with Security Clearance
NewBase-2 Solutions, LLC
Bethesda, MD🇺🇸$10k/yrHybridYesterdaySAFeMicroservicesLogstash+17Technology - BS
DevOps Software Engineer with Security Clearance
NewBase-2 Solutions, LLC
Bethesda, MD🇺🇸$10k/yrHybridYesterdayMicroservicesLogstashAgile+12Technology - BA
Senior DevOps Engineer
NewBooz Allen Hamilton
Washington, DC🇺🇸$99k - $225k/yrOn-site17 hours agoDockerShellAWS+4Technology - JM
Vice President - Senior Manager of Site Reliability Engineering (Chief Data & Analytics Office)
NewJ.P. Morgan
Jersey City, New Jersey🇺🇸On-site14 hours agoDockerAWSSplunk+8Technology