Why This Role Stands Out
This Machine Learning Engineer role offers a unique opportunity to build and own critical ML systems that drive autonomous research, making you instrumental in accelerating scientific discovery. If you thrive on bridging the gap between cutting-edge research and robust production infrastructure, this is your chance to make a significant impact and grow your expertise. Apply today to join a dynamic team at the forefront of AI innovation.
Quick Overview
Job Description
ABOUT THE COMPANY
We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site
ABOUT THE ROLE
You'll build and maintain the ML systems and pipelines that our research runs on top of data pipelines, training infrastructure, evaluation tooling, deployment, observability. The work bridges research and production, and you'll be the person who makes "we ran an experiment" actually mean "we ran it correctly, at scale, with results we trust."
This is a senior ML engineering role. You'll own systems end-to-end. You'll work with researchers daily and translate research code into infrastructure that the team can rely on. You'll move fast and you'll be measured on whether your systems make the team faster.
WHAT YOU'LL DO
- Build and maintain the training, evaluation, and deployment pipelines that our research runs on
- Take research code from prototype to production: refactor, harden, instrument, test
- Design observability into our ML systems (metrics, logs, traces, eval dashboards) so failures surface fast
- Own data pipelines for training and evaluation: ingest, dedup, version, validate
- Work closely with researchers to understand what they need, what's slow, and what's brittle
- Set engineering standards across our ML stack (testing, reviews, runbooks) so the team scales
- Contribute to architectural decisions that shape how research and
production interacts
WHAT WE'RE LOOKING FOR:
- Senior ML engineer with 6+ years building production-grade ML systems
- Track record across the full lifecycle: data, training, evaluation, deployment, monitoring
- Strong distributed systems experience; you've shipped systems that have to
be on
- Fluent Python, fluent with at least one of (PyTorch, JAX); comfortable at the systems-level when needed
- Comfortable with experimentation infrastructure (Ray, Slurm, Kubernetes, or
similar)
- Bias toward shipping; you prefer working code over working diagrams
- Strong written communication
NICE TO HAVE:
- Experience building experimentation platforms or research infrastructure
at a frontier ML lab
- Background in distributed training systems
- Open-source contributions to ML infrastructure
- History of working effectively with small senior teams
THIS ROLE IS PROBABLY NOT FOR YOU IF:
- You want to do research with engineering as a side activity: this is engineering as the main thing
- Cross-functional work with researchers (translation, scoping, education) doesn't appeal
- Long-running ownership of running systems isn't appealing: this role has it
Similar jobs
- PI
Staff Machine Learning Engineer, Content Visual AI
NewAuto ApplyPinterest
San Francisco🇺🇸Remote5 hours agoMachine LearningHiveLLMTechnology - HE
Inference Optimization Engineer
NewAuto ApplyHedra
San Francisco🇺🇸Hybrid7 hours agoCUDADeep LearningC+++2Engineering - ST
Machine Learning Engineer, Radar
NewAuto ApplyStripe
Seattle🇺🇸Hybrid22 hours agoSQLDeep LearningPyTorch+1Technology - TR
Machine Learning Engineer, Applied
NewAuto ApplyTracelabs
United States🇺🇸Remote22 hours agoRoboticsComputer VisionDeep Learning+3Technology - CO
Staff Machine Learning Engineer, Personalization
NewAuto ApplyCoupang
Mountain View🇺🇸$152k - $277k/yrHybridYesterdayAWSMLflowMachine Learning+11Technology - RM
Member of Technical Staff, ML Platform
NewAuto ApplyRunway Ml
Remote🇺🇸RemoteYesterdayRoboticsKubernetesLESS+3 - DE
Staff Inference Engineer
NewAuto ApplyDesignworkstalent
Bellevue🇺🇸HybridYesterdayEngineering - BU
Staff Machine Learning / Operations Research Engineer
NewAuto ApplyBurq, Inc.
United States🇺🇸Remote23 hours agoMLOpsDeep LearningForecasting+2Technology - OP
Senior AI/ML Test and Evaluation Engineer
NewAuto ApplyOpenTeams
United States - Remote OR Hybrid🇺🇸$145k - $250k/yrRemoteYesterdayExpressMachine LearningNumPy+6Technology - CL
Machine Learning Engineer (AI/ML)
NewAuto ApplyClickhouse
New York🇺🇸RemoteYesterdayGCPAWSMachine Learning+5Technology - WO
AI/ML Engineer
NewAuto ApplyWoolpert
Remote - United States🇺🇸$118.2k - $147.8k/yrRemote2 days agoGCPAgileArticulate+7Technology - PI
Director, Machine Learning Engineering, Ads Quality
NewAuto ApplyPinterest
Palo Alto🇺🇸$314.6k - $550.5k/yrHybrid2 days agoMachine LearningTechnology