Haystack
← Back to Jobs
Technology
AT

Applied AI Engineer (Agent Skills)

Advanced Tech PlacementJohns Creek, GA🇺🇸United StatesPosted Sep 30, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Johns Creek, GA, United States
Posted
20 hours ago
LLM

Job Description

We are looking for a AI Evaluation Engineer (Agent Skills)

This role focuses on building reusable AI skills that support the component factory and improve how teams implement applications with our design system. The position centers on shared capabilities such as requirements gathering, planning, accessibility review, spec generation, and implementation support using design system components and utilities.

Responsibilities:

  • Build reusable core skills that support the factory workflow, including requirements gathering, planning, accessibility review, and spec generation.
  • Create AI-assisted implementation skills that help teams use design system components and utilities correctly in application code.
  • Improve how shared design system guidance gets translated into practical implementation workflows.
  • Partner with engineering and design to identify repetitive implementation problems that can be improved through reusable skills.
  • Help define skills that increase consistency, reduce rework, and make application delivery faster within design system boundaries.
  • Contribute to a shared AI capability model that supports both upstream definition work and downstream implementation work.

Requirements:

  • Strong software engineering fundamentals.
  • Hands-on experience building applied AI capabilities, workflow tools, or agent-like systems.
  • Ability to create reusable skills that support both factory workflows and application implementation.
  • Experience working with component libraries, shared utilities, or design system-based application development.
  • Strong understanding of how engineering teams move from requirements and specs into implementation.
  • Strong judgment about how to balance standardization, usability, and product-team flexibility.
  • Ability to work across ambiguity and turn broad ideas into practical, usable internal capabilities.

Required Skills:

  • An engineering foundation; it's an engineer who evaluates, not a pure QA person.
  • Deep understanding of web standards and how the web fundamentally works to judge whether generated component code is actually right.
  • Experience writing and evaluating agent skills and prompts, including knowing which prompt patterns work better than others.
  • Experience testing and evaluating AI and LLM output, including eval frameworks, rubrics, LLM-as-judge approaches, regression suites for prompts, and similar work.
  • Proficiency in JS/TS, with the same maintainability consideration as Role 1.

Preferred Skills:

  • Write agent skills and evaluate their quality so the instructions and prompts inside skills reliably get the best results from AI agents.
  • Build evaluation mechanisms for output that is non-deterministic, where you can't just check for an exact match.
  • Handle a second source of change: skills draw on the team's own documentation and component APIs, which keep evolving. The system must catch regressions when either the skills or the docs change.
  • Expand what they've started with "golden prompts." These are prompts re-run after every change to confirm output is still correct. The open problem is how to define and measure "correct" when output is never identical twice.
  • Ensure what product teams receive keeps getting better, not worse, as changes ship.

Similar jobs