Quick Overview
Job Description
The mission
Nolla is building AI doctors so that top-quality healthcare is accessible to everyone. Nolla Derm is the #1 medical skincare treatment app in the U.S. App Store. We've treated thousands of acne patients in the U.S., scanned 1% of Norway's population for skin cancer, and our in-house clinical models are state of the art on clinical benchmarks. We just launched NollaMD, our urgent care app, and we're building dedicated specialty apps for conditions like women’s health. We’ve raised $6.5M from General Catalyst and other strategic investors.
The role
We have built our own clinical benchmarks, post-trained on them, and produced models that beat frontier performance on clinical tasks. We are in a unique position of both delivering care to real patients and training the models that deliver it. You’ll work on our evals and harnesses, build the data loop with clinicians that feeds them, and lead where they go next.
You'll work directly with the founding team and with our post-training partners, and your work reaches patients immediately. This is a hands-on role: you will write code most days, design studies, and be the author of record on what we publish.
What you'll build
Evals and benchmarks
Design and maintain the benchmarks we use to judge clinical accuracy, safety, and documentation quality
Extend them to our agentic system: tool use, long-horizon tasks that span many visits, memory of a patient's history, uploaded records, labs, and escalation behavior
Graders and harnesses
Build the grading stack: deterministic checks, rubric graders, and LLM judges calibrated against clinician ratings
Improve the production harness the product runs on
Data loop with clinicians
Turn real, consented cases into eval cases and training examples, with clinician review, de-identification, and provenance built in
Run the clinician review workflows that produce rubrics, labels, and feedback at scale
Research
Lead our clinical evaluations from protocol through publication, and publish our benchmarks for the field
Collaborate with post-training on reward design and data curation, and run experiments where it helps
What you bring
You have built eval infrastructure for LLM systems that other people depended on
You have opinions about contamination, difficulty calibration, judge bias, and reward hacking
You can write a paper. Authorship on empirical ML work, ideally with a human comparison or a benchmark release
You are excited to work with some of the best physicians in the country and turn their judgment into model performance
You've worked on small teams and built things from zero
Bonus points
Clinical AI evaluation experience: rubric-based health benchmarks, simulated-patient studies, agentic clinical benchmarks
Post-training experience (RLHF, DPO, GRPO or similar)
Experience with multimodal models
Familiarity with eval and environment frameworks such as Inspect, Verifiers, or Harbor
Role logistics, compensation & benefits
Role Type: Engineering
Salary: $180,000–$240,000, based on experience
Job Type: Full-time
Work Setup: In-person, New York City
Equity: Meaningful equity, commensurate with experience
Health Insurance: Medical, dental, and vision
HSA/FSA: Eligible
Time Off: Flexible, unlimited vacation
Additional Perks: Meal stipends, team retreats
Work Hours: Flexible but demanding. We're building something that matters
Growth: Founding team members step into expanded roles as we scale