Why This Role Stands Out
This remote role offers a unique opportunity to build and own reliability from the ground up at a rapidly growing AI company, with the chance to ship code and define critical infrastructure. You'll thrive here if you're passionate about ensuring 99.99% uptime, have a knack for incident command, and are eager to contribute to a robust engineering culture. Apply now to make a significant impact on a platform powering millions of calls!
Quick Overview
Job Description
Vapi (/ˈVɑːpi/):
Voice AI that resolves, not transfers
Powering 1 billion calls for companies like Amazon Ring, Intuit, ServiceTitan, and New York Life
Trusted by 1 million developers building the future of voice agents
Backed by Peak XV, Bessemer, Kleiner Perkins, M12, Y Combinator, and more with $72M raised
Why We’re Hiring This Role:
99.99% call completion is the number this role drives. Vapi runs live phone calls — a p99 spike means callers drop. We’ve had 15 stability-gap outages worth learning from, and we need someone who runs incident command, owns SLOs and error budgets, and builds the reliability culture from scratch.
This is not a bash-and-YAML role. You’ll ship code (Go or TypeScript) for services that monitor and manage the platform: auto-remediation, capacity forecasters, oncall tooling. Capacity planning, load testing, and KEDA-based autoscaling for Vapi’s wscaler and workerpool-cron-scaler are on your plate.
What You’ll Do:
30 Day: Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path.
60 Day: Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the highest-impact services. Run the first proper load test against provider rate limits and per-org concurrency. Tune autoscaling for wscaler / workerpool-cron-scaler.
90 Day: Ship a real platform service — capacity forecaster, auto-remediation, or oncall tooling — in Go or TypeScript. Own the postmortem process. Drive a measurable improvement in p99 call completion or MTTR.
Who You Are:
Must-haves
You’ve run incident command and postmortem discipline at scale on a real oncall rotation.
You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog.
You’ve done capacity planning and load testing for production systems with real users.
You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown.
You know backpressure and autoscaling patterns — KEDA, custom metrics scaling.
Nice-to-haves
You ship code, not just scripts. You can build platform services in Go or TypeScript (matches Vapi’s cluster-manager, database-health, wscaler, incidentManager).
Real-time / latency-sensitive product background where degraded means a dropped call, not a slow dashboard.
Tech stack you’ll work in
Languages: Go and TypeScript (you ship code, not just scripts), Bash.
Observability: Chronosphere, Prometheus, Grafana, Datadog, OpenTelemetry.
Orchestration: Kubernetes on EKS — production ops (HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown, pod crash diagnosis).
Autoscaling and backpressure: KEDA, custom metrics scaling (matches Vapi’s wscaler and workerpool-cron-scaler).
Load testing: script-based load testing, provider rate-limit auditing, per-org concurrency auditing.
Vapi services you’ll touch or build: cluster-manager, database-health, wscaler, incidentManager.
Where you likely come from
A real-time / latency-sensitive product (Discord, Zoom, Mux, Twitch, Twilio, LiveKit, Cloudflare, a trading firm, a gaming backend), or a FAANG SRE / Production Engineer (Google, Uber, Twitter/X, Meta) who misses being hands-on.
Weak fit: SRE from analytics or CRM backends where “degraded” means a slow dashboard, not a dropped call. Anyone uncomfortable reading or writing code.
Why Vapi:
Generational impact: Build the human interface for every business
Ownership culture: 70% of the company are previous founders
Kind team: The founders, Jordan and Nikhil, are Canadians
Tier-1 Investors: YC, KP seed, Bessemer Series A
What We Offer:
Real stake: We offer a competitive salary and excellent equity ownership
Comprehensive health coverage: medical, dental, and vision plans
Team love: We love hanging out, and we do quarterly off-sites
Flexible time off: take what you need
More: catered meals, transportation, gym, and a $10k annual L&D budget
Similar jobs
- VS
SRE Director || ONSITE - Hybrid || Phoenix, AZ
NewValue Spectrum Technologies LLC
Phoenix, AZ🇺🇸HybridYesterdayMicroservicesKubernetesStakeholder ManagementTechnology - SP
Site Reliability Engineer (Manufacturing Infrastructure)
NewSpaceX
Bastrop, TX🇺🇸Hybrid14 hours agoDockerAnsibleKubernetes+2Technology - AS
Linux Engineer
NewApex Systems
Providence, RI🇺🇸Hybrid14 hours ago401kComplianceRoot Cause Analysis+1 - OR
Senior Site Reliability Engineer
Oracle Corporation
Nashville, TN🇺🇸$81.1k - $187k/yrHybrid3 weeks agoOracleAnsibleBash+4Technology - OR
Lead Principal Site Reliability Engineer - (work in Vienna VA location)
Oracle Corporation
Vienna, VA🇺🇸$96.3k - $264.1k/yrHybrid3 weeks agoDockerOracleAWS+20Technology - ES
Senior Data Platform/DevOps Engineer
NewEPAM Systems
United States🇺🇸Hybrid14 hours agoMongoDBNeo4jEncryption+14Technology