Quick Overview
Job Description
About the Role
Build and operate the infrastructure that supports research and engineering teams working on biomedical AI. As an early hire, you will establish the platform foundations for model training, inference, and data workflows, while shaping how the systems scale.
What You'll Do
Design, build, and operate cloud and GPU infrastructure for training and serving large models.
Develop platform capabilities for orchestration, job scheduling, CI/CD, developer environments, and internal tooling.
Own reliability, observability, security, and cost optimization across compute and data systems.
Partner with researchers and ML engineers to identify and resolve infrastructure bottlenecks.
Establish infrastructure-as-code practices and operational standards.
What We're Looking For
5 to 10+ years of experience in platform engineering, infrastructure, or site reliability engineering, including experience in a high-growth startup or strong engineering organization.
Hands-on experience with AWS or GCP, Kubernetes, and infrastructure-as-code tools such as Terraform.
Strong programming skills in Python and/or Go.
Experience operating ML infrastructure, including GPU clusters, distributed training, and large-scale data pipelines.
Comfort taking ownership and making technical decisions in an early-stage environment.
Compensation & Benefits
Visa sponsorship is not available for this role.
Location
On-site in San Francisco, California, United States. Candidates should be based in or willing to relocate to San Francisco.
Similar jobs
- MS
Senior DevSecOps Engineer
NewAuto ApplyMomentus Space LLC
San Jose🇺🇸Hybrid5 hours agoShellAWSRobotics+11Engineering - LI
DevOps Engineer
NewAuto ApplyLIGHTFEATHER IO LLC
United States🇺🇸Hybrid8 hours agoDockerRubyAWS+4Technology - LI
Staff Software Engineer (DevOps / Platform Engineering)
NewAuto ApplyLiberate
Boston / San Francisco🇺🇸Hybrid10 hours agoAWSComplianceGitHub Actions+8Technology - FO
DevSecOps Engineer - Leesburg
NewAuto ApplyFortreum
Leesburg🇺🇸Hybrid9 hours agoGCPAWSSplunk+13Engineering - FO
DevSecOps Engineer - Reston
NewAuto ApplyFortreum
Reston🇺🇸Hybrid9 hours agoGCPAWSSplunk+13Engineering - EN
Senior DevOps Engineer
NewAuto ApplyEncoura
Remote🇺🇸Remote12 hours agoDockerMongoDBSQL+14Technology - TA
Staff Infrastructure Engineer
NewAuto ApplyTabs
New York City🇺🇸On-site8 hours agoDockerAWSCompliance+8Technology - RE
Platform Engineer
NewAuto ApplyResend
Americas🇺🇸Remote10 hours agoNode.jsAWSLoad Balancing+6Technology - 2K
Senior Site Reliability Engineer
NewAuto Apply2K
Austin🇺🇸Hybrid9 hours agoMySQLPackerAWS+18Technology - OP
Site Reliability Engineer, Provider Operations
NewAuto ApplyOpenrouter
Remote (US)🇺🇸Remote8 hours agoGCPAPI GatewayLoad Balancing+7Technology - MI
Senior DevOps Engineer – PostgreSQL & Kafka on Kubernetes
NewAuto ApplyMirantis
Remote🇺🇸Remote10 hours agoAWSMachine LearningSOC 2+15Technology - TH
Senior DevOps Engineer
NewAuto ApplyTheIncLab
Colorado Springs, Colorado🇺🇸HybridYesterdayAWSAnsibleBash+12Technology