Haystack
← Back to Jobs
Full time
Other
CH

Research Engineer - ML Infrastructure

ChaidiscoverySan Francisco office🇺🇸United StatesPosted Sep 24, 2026

Quick Overview

Seniority
Mid Senior
Employment type
Full Time
Work mode
Hybrid
Location
San Francisco office, United States
Posted
21 hours ago
CUDAPyTorchPython

Job Description

About Chai Discovery

Chai builds the design suite for molecules. We train frontier models that learn the underlying foundations of biochemical structure and interaction, so scientists can move faster and pursue targets that other methods cannot reach.

AI is reinventing life sciences the same way it reinvented software engineering, and Chai is at the forefront of this shift. Leading pharmaceutical companies like Eli Lilly, Pfizer, and Novartis are adopting our platform to power their drug discovery programs.

We value diverse perspectives and are ready to find greatness in unexpected places.

About the role

Research Engineers on ML Infra make our models train and run performantly, reliably, and at scale by owning the distributed systems that sit underneath every model our researchers ship. As a ML Infra Research Engineer, you will:

  • Architect, debug, and optimize the distributed ML training stack across the model, layer, and kernel levels — eliminating runtime and reliability bottlenecks on large GPU clusters.

  • Profile end-to-end training runs to find bottlenecks across compute, communication, and storage, and build tooling to monitor throughput, utilization, and uptime across clusters.

  • Optimize ML workloads through parallelism strategies, quantization, and custom CUDA/Triton kernels.

  • Work closely with Research Scientists to ensure new model architectures and training recipes scale efficiently, from early experiments to frontier-scale runs.

  • Own reliability of the training stack: fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.

About you

  • 4+ years of industry experience working within AI/ML infrastructure teams.

  • Proficiency in Python and PyTorch or JAX.

  • Strong software systems design skills, with comfort operating across the stack from model code down to kernels.

  • Experience with orchestrating GPU clusters and large-scale model training.

  • Experience with optimizing ML workloads: parallelism, quantization, CUDA/Triton kernels.

We offer

The opportunity to work at the vanguard of AI research and frontier biology, with world-class people, on a mission that matters. We protect & promote a culture of high velocity and ownership. We compensate our team accordingly.