Site Reliability Engineer - Onsite at San Francisco or Sunnyvale, CA
Quick Overview
Job Description
Core Production Engineer - Onsite in San Francisco or Sunnyvale, CA.
What You'll Be Working On:
As a CORE PE, you will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention
You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day.
Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
You'll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment
Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
By day's end, you will document your work, share insights with your team, and plan for the next day's challenges, always with a customer-centric mindset.
What You’ll Bring to the Team:
Strong experience with architecture, design patterns, reliability and scaling of new and current systems
Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions
Experience building observability from the ground up — defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do
Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem
Experience writing high quality code with at least one programming language (Python, Go, or similar)
Experience with system-level debugging, including kdump, and kernel panic analysis.
Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.
Experience with TCP/IP and network programming
Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms.
Hardware and GPU troubleshooting experience (nice to have)
Exposure to OVN/OVS-based networking stack (nice to have)
Skills
Similar jobs
DevOps I – Linux, AWS, Release Management
Partner's Consulting, Inc. · Englewood, United States
2 minutes agoReferral Network Engineer - DHA NE&S with Security Clearance
ASRC Federal · United States
5 minutes agoAWS DevOps Engineer
ASCII Group LLC · McLean, United States
11 minutes ago$57/hrDevOps Engineer
Genesis10 · Charlotte, United States
12 minutes agoSenior Site Reliability Engineer with FedRAMP - Remote
Momento USA LLC · United States
28 minutes agoPSOC DevOps Engineer
Genesis10 · Dallas, United States
28 minutes ago