Haystack
← Back to Jobs
Technology
HT

Software Engineer — AI Evaluation & Automation

HonorVet TechnologiesUnited States🇺🇸United StatesPosted 11 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
DockerGitJavaJavaScriptLLMPython

Job Description

HonorVet Technologies is a Service Disable Veteran-Owned IT staffing firm, ISO 9001, and ISO 27001 certified, working with federal agencies, state governments, and Fortune 500 enterprise clients across the US. What makes us different isn't a tagline; it's the way we work. We don't forward resumes and hope for the best. We take the time to understand where a professional like you are headed and only reach out when we genuinely believe there's a fit worth exploring. I came across your profile, and something stood out.

Title: Software Engineer — AI Evaluation & Automation
Duration: 06 months to start with the high possibility of extension
Work Schedule: Remote

Role Overview: Help build and scale the tooling we use to measure how well AI-powered software development tools perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.

Key Responsibilities

  • Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated work trees) so results stay comparable over time.
  • Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
  • Support execution-based benchmarking across quality, productivity, and efficiency measures, including cost and latency.
  • Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
  • Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.

Required Skills & Experience

  • Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
  • Proficient in at least one general-purpose language such as Python, Java, or JavaScript — the specific language background is flexible.
  • Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
  • Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
  • Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how
 benchmark contamination happens.
  • Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
  • Hands-on experience using AI coding tools and agentic harnesses such as Claude Code, Devin, or Cursor, and command of the best practices for working with them effectively.

Preferred Experience

  • Experience designing benchmarks or evaluations for software systems, especially execution-based grading that verifies against tests.
  • Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them against human raters.
  • Experience with build-system-aware test selection, such as Bazel or mapping changed files to the tests that cover them.
  • Experience building reproducible test environments and managing versioned evaluation datasets.
  • Comfortable writing up methodology and results for engineering leadership.

Notes from the Manager

They must have experience using existing AI tools with harnesses such as Claude, Devin, Cursor etc. and the best practices using them. Evaluation and benchmarking experience (getting the measurement right, not just automating it). Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
The intention is to renew every 6 months, and if FTE positions become available contractors will be considered first
Must work Pacific Time zone. Not accepting candidates with more than a two-hour difference.

If you are interested, feel free to reach out to me directly at [].

Similar jobs