Haystack
← Back to Jobs
Full time
Engineering
CL

Research Engineer, Synthetic Data

CleraSan Francisco🇺🇸United StatesPosted 27 Aug 2026

Quick Overview

Seniority
Mid Senior
Employment type
Full Time
Location
San Francisco, United States
Posted
19 hours ago
DockerLLMPython

Job Description

About the Role

This is a Research Engineer role focused on building synthetic data pipelines for AI agent training, sitting within a ~15-person engineering team of Olympiad medalists and published researchers. You'll design generation methods, validation systems, and quality metrics that directly expand model capabilities — work that sits at the frontier of RL-based AI alignment.

What You'll Do

  • Build end-to-end synthetic data pipelines that transform domain-specific workflows into structured, challenging training tasks for AI agents.

  • Collaborate with subject-matter experts to develop synthetic tasks spanning professional and technical domains.

  • Design task generation methods that produce diverse, realistic, and learnable training examples.

  • Build tooling to mutate, validate, and iteratively improve synthetic task quality.

  • Analyze model and agent performance on synthetic tasks to understand learning outcomes and failure modes.

  • Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.

What We're Looking For

  • 2–4 years of experience in software engineering, ML engineering, or AI research roles delivering data pipelines, ML infrastructure, or synthetic data systems.

  • Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.

  • Proficiency in Python and experience developing in Linux environments using containerization tools such as Docker.

  • Demonstrated understanding of synthetic data quality criteria and evaluation metrics — diversity, realism, learnability — and their inherent limitations.

  • Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.

  • Experience building automated systems to generate, validate, mutate, or process structured datasets at scale.

  • Track record of independently owning and delivering technical projects end-to-end with minimal predefined requirements.

  • Ability to detect edge cases, inconsistencies, and quality issues in synthetic or algorithmically generated datasets.

  • Comfort operating in unstructured, early-stage environments and reasoning from first principles.

  • Strong communication skills for effective remote collaboration across time zones.

  • Nice to have: Familiarity with reinforcement learning paradigms, agentic AI workflows, or LLM post-training pipelines.

Compensation & Benefits

Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available.

Location

On-site in San Francisco, CA, United States. Singapore-based candidates are also considered.

Similar jobs