Haystack
← Back to Jobs
Remote
Technology
OP

Senior AI/ML Test and Evaluation Engineer

OpenTeamsUnited States - Remote OR Hybrid🇺🇸United StatesPosted Oct 7, 2026

Why This Role Stands Out

This remote role offers a highly competitive salary and the chance to build the core evaluation capabilities for an innovative AI platform, ideal for a meticulous engineer passionate about refining AI performance. You'll thrive here if you possess a keen eye for detail and a desire to develop robust testing methodologies in a dynamic, forward-thinking environment. Apply today to make a significant impact in the evolving AI landscape!

Quick Overview

Salary
$145k - $250k/yr
Seniority
Mid Senior
Work mode
Remote
Location
United States - Remote OR Hybrid, United States
Posted
5 hours ago
ExpressMachine LearningNumPySciPyComputer VisionHugging FaceJupyterPyTorchPython

Job Description

Who We Are

Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.
OpenTeams exists to make ownership possible.
Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.
If that sounds like your kind of work, we'd like to meet you.

Senior AI/ML Test and Evaluation Engineer

Location: U.S - Remote OR Hybrid - Washington, DC, Denver, CO or Colorado Springs, CO. 

Work Authorization: U.S. citizenship required

Clearance:  U.S.-Remote Opening: An active clearance is not required. Candidates must be eligible and willing to obtain and maintain a U.S. security clearance. Hybrid Opening: An active TS/SCI clearance is preferred. Candidates may also be considered if they previously held a TS/SCI with CI polygraph or currently hold an active TS or Secret clearance.

Salary Range: $145,000–$250,000 USD, dependent on experience level and location

Openings: Two positions are available:

  • One hybrid position: Candidates must be located in or willing to work hybrid from Washington, DC; Denver, CO; or Colorado Springs, CO. An active TS/SCI clearance is required. Up to 15% travel is required.
  • One U.S.-remote position: Candidates may work remotely from anywhere in the United States. An active clearance is not required, but candidates must be willing and able to undergo the process required to obtain and maintain a U.S. security clearance.

About the Role

We're looking for a Senior AI/ML Test and Evaluation Engineer to build and operate the benchmarking and evaluation capability at the core of an AI platform. This is a role for someone who is more interested in what a model gets wrong than in what it gets right.

You build the evaluation harnesses — automated metrics paired with structured human expert judgment, applied to candidate models and to the agentic workflows built on top of them. You develop repeatable methodologies for comparing performance against current operational baselines, which means the comparison holds up when someone runs it again in six months with a different model. And you document the limitations and surface the failure modes that matter, including the ones nobody asked about.

Your reports go to senior stakeholders and inform decisions about which capabilities are ready to field. That's the weight of the job: a benchmark that looks good and hides a failure mode is worse than no benchmark at all, and you're the check against that.

This is hands-on engineering on an open-source toolchain.

Key Responsibilities

  • Design and implement a platform for test & evaluation of AI models and Agentic systems
  • Develop evaluation methodologies combining human and AI expert judging and multiple input and output modalities
  • Design and enhance user interfaces such that platform customers can express workflows in an intuitive and generic manner
  • Collaborate closely with subject matter experts, Model/Agent developers, and evaluation designers to ensure coverage over existing and to be discovered requirements.
  • Architect and implement robust experiment provenance and result tracking systems.
  • Ensure performance across integrated tools and scaling systems.
  • Provide technical mentorship, code review, design guidance, and clear documentation to support team knowledge transfer and user onboarding.
  •  

Required Skills & Experience

  • Proven experience building and deploying software platforms with complex integration surface, preferably handling Machine Learning or agentic workloads.
  • Deep experience using python, including API frameworks and agentic harnesses.
  • High level knowledge of Computer vision, language modeling, and/or task specific agent evaluation.
  • Hands on experience with common frameworks and tooling, eg PyTorch and Hugging Face ecosystem.
  • Strong communication skills and experience collaborating effectively with cross functional teams including external stakeholders.
  • Demonstrated experience or potential for leadership, especially in making architectural decisions, coordinating the efforts of other engineers, and developing standards and best practices.

Nice to Have

  • Currently hold or eligible to hold U.S. Security Clearance (Secret or higher)
  • Prior experience designing test and evaluation systems for AI
  • Experience deploying complex software systems in IL 4 or higher environments.
  • Experience participating in security assessment and authorization, hardening systems, and remediating vulnerabilities. 
  • Previous experience contributing to open source projects and participating in open source communities.
  • Exposure to the intelligence community or department of defence subject matter experts. 

Grow With Us

At OpenTeams, growth isn’t just about the company—it’s about you.
We believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.

Opportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily.  That global perspective and diversity makes our solution more universal and robust.  We are committed to continuing to celebrate diversity on our team.

Supported people are successful people.  We offer 100% employer paid medical premiums for employees and self-managed PTO with a minimum time off requirement, so that our teams are able to do their best work.
We invest  in curiosity, creativity, and ownership. That means you’ll be trusted to boldly innovate, supported to learn fast, and celebrated for successful collaboration.

Commitment to diversity, equity, inclusion, and belonging

OpenTeams understands that valuing diverse creative practices and forms of knowledge is crucial to and enriches the company’s core mission. We encourage applications from everyone, including members of all equity-seeking communities, such as (but certainly not limited to) women, racialized and Indigenous persons, disabled people, persons of all sexual orientations, gender identities and expressions.

We are an equal opportunity employer - all qualified applicants will receive equal consideration for recruitment, interviews, employment, training, compensation, promotion, and related activities. We do not discriminate based on race, religion, gender, gender identity, gender expression, color, national origin, pregnancy, ancestry, domestic partner status, disability, sexual orientation, age, genetic predisposition, medical condition, marital status, citizenship status, military or veteran status, or any other basis covered by applicable laws. OpenTeams will not tolerate discrimination or harassment based on these characteristics or any other unlawful behavior, conduct, or purpose.

Similar jobs