Haystack
← Back to Jobs
Other
IN

Engagement Manager, Agentic AI Workflow Evaluations

Innodata Inc.In Office - San Jose🇺🇸United StatesPosted Sep 19, 2026

Why This Role Stands Out

This role offers a unique opportunity to lead a cutting-edge AI evaluation team and directly contribute to the responsible advancement of artificial intelligence with a globally recognized data engineering company. You'll thrive here if you possess strong leadership skills, a keen eye for detail, and a passion for AI, making a significant impact on a frontier AI project. Apply now to be at the forefront of AI innovation and professional growth.

Quick Overview

Seniority
Mid Senior
Location
In Office - San Jose, United States
Posted
14 hours ago
Generative AIOnboardingPerformance Management

Job Description

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

We are standing up a dedicated onsite team to evaluate complex, real-world agentic AI workflows for a frontier AI customer. The Engagement Manager owns this team and this engagement: a group of ten reviewers and one QA lead working inside isolated test environments, assessing whether AI agents complete tasks safely, respect user intent and consent, and hold up under close scrutiny. 

This is a delivery leadership role with a strong customer-facing component. You are the customer's primary day-to-day contact at the site, accountable for throughput, quality, and turnaround, and responsible for identifying where the engagement should grow next. You will not be scoring trajectories as a daily task, but you need enough technical depth to audit the team's work yourself and defend a scoring decision in a room with the customer's technical leads. 

What You’ll Own:

  • Own end-to-end delivery for the onsite team: throughput, quality, turnaround, and capacity against committed volumes 
  • Manage and develop a team of eleven, including hiring, onboarding, performance management, and retention 
  • Serve as the primary onsite point of contact for the customer's program and technical leads; run recurring reviews, report on delivery metrics, and resolve issues before they escalate 
  • Audit reviewer output directly — sample trajectories, check rubric application, and assess whether scoring rationale would survive customer review 
  • Partner with the QA lead on calibration cycles, quality metrics and evaluation, and rubric refinement; arbitrate disagreements that calibration does not resolve 
  • Own the escalation path for safety-relevant and ambiguous findings, including judgment on what reaches the customer and how quickly 
  • Identify expansion opportunities within the account: adjacent workstreams, new task types, additional capacity, and scope the work with internal delivery and commercial teams 
  • Forecast staffing and cost against the engagement's commercial model, and flag variance early 
  • Maintain information security, privacy, and facility access practices required by the customer's onsite environment 

You’ll Thrive in This Role If You Have:

  • Bachelor's degree or equivalent practical experience 
  • 5+ years of professional experience, including direct management of a delivery team of 10–20 people 
  • Track record as the customer-facing owner of a services or delivery engagement with an enterprise or technology client, including metrics reporting and issue escalation 
  • Significant experience with AI/ML evaluation work: annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review 
  • Ability to read rubric-scored work critically and form an independent view of whether a score is correct 
  • Strong written communication; able to produce reporting and escalation memos that a technical customer will accept without rework 

The expected hourly salary range for this position is $75-85 p/hour, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams. 

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Similar jobs