Why This Role Stands Out
As a Senior Data Engineer at HealthLeap, you'll build and operate critical data pipelines that directly impact patient care outcomes and drive significant company growth, all within a hybrid work environment that supports flexibility. This role is ideal for an engineer passionate about transforming complex clinical data into reliable, usable formats for AI and analytics, offering substantial opportunities for skill development and career advancement within a rapidly expanding, mission-driven organization. You'll thrive here if you're eager to contribute to a groundbreaking AI operating system that's revolutionizing healthcare and want to be part of a dynamic team making a real difference.
Quick Overview
Job Description
About Healthleap
Every day, millions of hospitalized patients who need intervention are missed because clinicians simply can't see everything. HealthLeap is building the AI operating system that helps care teams identify these missed patients, enabling them to improve health outcomes and generate millions of dollars. HealthLeap is changing what is possible: closing gaps that traditional workflows and clinician capacity could never.
Over the past year, we've grown contracted revenue more than 13x, expanded rapidly across leading health systems, and now help care teams identify patients across millions of inpatient encounters.
We're ~25 people. >$32M raised. SF-based, hybrid-friendly. And, we're delivering results that are changing lives.
Senior Data Engineer
About the role
HealthLeap runs on data. We’re live at 40+ hospitals and plan to add another 100. You’ll build the core pipelines that move messy clinical data from hospital systems into model inputs, analytics, and the APIs behind what clinicians see.
You’ll make that data usable and reliable. When a feed arrives late, a field changes, or a number looks wrong, you’ll trace it through the system and fix the cause.
What you’ll do
Build and operate pipelines from hospital ingestion through transformation and delivery.
Produce trusted data for ML pipelines, customer analytics, and user-facing APIs.
Define data contracts and checks that catch missing records, schema changes, and incorrect values.
Handle backfills, late data, failures, and recovery.
Work with integration, ML, and product engineers to get data reliably where it needs to go.
What we’re looking for
5+ years building production data systems, with strong Python and SQL.
Experience owning pipelines that depend on messy, changing external data.
Strong data modeling and judgment about correctness, monitoring, and recovery.
The ability to trace a problem across systems and own the fix through production.
What will make you stand out
Experience building data pipelines for ML products.
Experience with clinical data, EHRs, HL7, or FHIR.
Early-stage experience building and operating core data systems.
This role is NOT for you if
You want to own one piece of the data stack. You’ll work across ingestion, model inputs, analytics, and product APIs.
You want predictable 9-to-5 hours. We protect deep rest, but a hospital go-live can mean a 60+ hour week.
Interview process
No LeetCode or puzzles. Use the tools you’d use on the job, including AI.
Intro call
Data pipeline design
Practical data exercise
Onsite with the team in San Francisco
We decide the same week as the onsite.
Compensation and benefits
$175,000–$275,000 base plus meaningful equity
100% covered healthcare premiums
Unlimited PTO with a 20-day minimum
4% 401(k) match
Laptop and home office budget
Location
San Francisco, in person. We work together in the office by default, with flexibility to work from home when needed. We judge output, not hours.
If you're passionate about applying frontier AI to real-world impact, join us in building healthcare's future.