Member of Technical Staff, Infrastructure (Kubernetes Specialist)
Quick Overview
Job Description
About TypeSafe
TypeSafe is an AI lab building intelligence beyond chat: AI designed to work inside software and power real-world automation. Our mission is to make intelligence composable—dependable enough for builders to invoke, inspect, constrain, and layer into production systems.
Today’s models are remarkably capable, but that intelligence is still hard to build on. We’re rethinking the AI stack from first principles so software can branch on meaning, judgment, and intent while code continues to handle exact computation.
We’re a small, fast-moving team from OpenAI, Google Brain, and Meta/FAIR. We care about technical rigor, real-world impact, and building systems people can trust. Join us to help turn today’s intelligence into the foundation for a new generation of software.
About the role
We're looking for an Infrastructure Engineer to build and operate the infrastructure behind TypeSafe AI's products at global scale. You'll own the systems that serve millions of users across regions — from provisioning Kubernetes clusters across multiple clouds to optimizing networking for low-latency AI inference.
This is a high-impact role on a small, fast-moving team. You'll work across the full infrastructure stack: cloud primitives, container orchestration, networking, observability, and the specialized infra that makes large-scale model inference efficient.
What you'll do
Design, deploy, and operate Kubernetes clusters across multiple regions and clouds
Build and maintain infrastructure for the platform that powers LLM inference workloads globally
Own networking, including VPCs, peering, load balancing, DNS, service mesh, CNI
Manage GPU infrastructure and autoscaling for ML workloads
Write and maintain infrastructure as code (Pulumi / Python)
Operate and improve observability: monitoring, alerting, tracing, logging
Requirements
Deep experience with Kubernetes in production at scale: networking, storage, scheduling, upgrades
Strong background in AWS
Hands-on experience with infrastructure as code (Pulumi, Terraform, or similar)
Solid understanding of Linux networking
Track record with high-traffic production ML systems
Programming fluency, Python preferred
Nice to have
Experience with large-scale LLM / ML inference infrastructure (GPU scheduling, model serving, vLLM, KubeRay, Kubernetes-native tooling)
Kubernetes networking depth with Cilium or other CNI plugins; service mesh (Istio, Envoy)
Multi-cloud infrastructure
Background in site reliability engineering including SLOs, incident response, capacity planning
Life at TypeSafe
We’re a small, flat, close-knit team working to make intelligence dependable enough to become part of everyday software. We work fully in person from our San Francisco office near Embarcadero station. We love what we do and care deeply about the work.
We strive for excellence and craftsmanship and won’t stop until we get there. When the team wins, we all win, and we enjoy collaborating and inspiring each other to grow—as a team and as individuals.
We value emotional honesty, kindness, and bringing your whole self to work. We build machines; we don’t try to be machines.
We want TypeSafe to be the place where you do the most impactful work of your career and help define our future as a company.
We provide
Base salary of $150k–250k plus equity, based on leveling
100% covered health insurance
Daily lunch and dinner
Visa sponsorships
401K plans
Similar jobs
- DW
DevSecOps Engineer
NewDark Wolf Solutions
Tampa, Florida🇺🇸$160k - $180k/yrHybrid48 minutes agoSAFeAgileEngineering - VS
SRE Director || ONSITE - Hybrid || Phoenix, AZ
NewValue Spectrum Technologies LLC
Phoenix, AZ🇺🇸HybridYesterdayMicroservicesKubernetesStakeholder ManagementTechnology - SP
Site Reliability Engineer (Manufacturing Infrastructure)
NewSpaceX
Bastrop, TX🇺🇸Hybrid15 hours agoDockerAnsibleKubernetes+2Technology - AS
Linux Engineer
NewApex Systems
Providence, RI🇺🇸Hybrid15 hours ago401kComplianceRoot Cause Analysis+1 - OR
Senior Site Reliability Engineer
Oracle Corporation
Nashville, TN🇺🇸$81.1k - $187k/yrHybrid3 weeks agoOracleAnsibleBash+4Technology - OR
Lead Principal Site Reliability Engineer - (work in Vienna VA location)
Oracle Corporation
Vienna, VA🇺🇸$96.3k - $264.1k/yrHybrid3 weeks agoDockerOracleAWS+20Technology