Why This Role Stands Out
This remote AI Evaluation Engineer role offers a unique opportunity to shape the quality and safety of cutting-edge AI systems, empowering you to build robust evaluation frameworks and directly impact product readiness. You'll thrive here if you possess strong analytical skills and a passion for ensuring AI excellence, making this an exciting step in your career. Apply today to contribute to a leading technology company and advance your expertise in a dynamic field.
Quick Overview
Seniority
Mid Senior
Work mode
Remote
Location
NJ, United States
Posted
1 week ago
LLM
Job Description
AI Evaluation Engineer
Location: USA Remote
Employment Type: Contract
Employment Type: Contract
Position Summary
We are seeking an AI Evaluation Engineer to establish and maintain the quality standards used to determine whether production AI and LLM systems are ready to launch.
This role will build centralized evaluation frameworks, create golden datasets, establish measurable quality thresholds, audit AI evaluation results, and implement continuous production monitoring. The successful candidate will bring strong technical evaluation expertise along with the independence and analytical rigor required to objectively determine whether AI systems meet production quality and safety standards.
Key Responsibilities
- Design and build centralized LLM/AI evaluation frameworks and reusable evaluation templates.
- Establish evaluation standards that can be adopted consistently across multiple AI engineering teams.
- Create and maintain golden datasets in partnership with business, domain, and subject-matter experts.
- Develop evaluation suites covering functional quality, accuracy, safety, reliability, and domain-specific requirements.
- Implement and assess LLM-as-a-judge evaluation approaches while accounting for their limitations and failure modes.
- Define measurable pass/fail thresholds and production-readiness criteria for AI applications.
- Independently review and audit evaluation suites developed by individual engineering teams.
- Build regression testing approaches for LLM applications, agents, prompts, retrieval systems, and model changes.
- Establish production monitoring for AI quality degradation, drift, regressions, and incidents.
- Analyze and report evaluation pass rates, quality trends, regressions, and production incidents.
- Partner with AI engineers, platform engineers, product teams, and domain experts to continuously improve AI quality.
Required Qualifications
- 4+ years of experience in ML/LLM evaluation, AI quality engineering, applied research engineering, or related areas.
- Hands-on experience designing and implementing AI/LLM evaluation frameworks.
- Experience developing golden/reference datasets and measurable evaluation criteria.
- Strong understanding of LLM-as-a-judge methodologies and associated failure modes.
- Experience evaluating LLM applications, agents, RAG systems, prompts, or other probabilistic AI systems.
- Strong statistical and analytical skills, including evaluation methodologies for relatively small sample sizes.
- Experience establishing quality thresholds and regression criteria.
- Strong engineering skills with the ability to build reusable evaluation tooling and automation.
- Ability to independently assess system quality and challenge release decisions when evaluation evidence does not meet established standards.
Preferred Qualifications
- Healthcare, clinical, or other safety-critical AI evaluation experience.
- Experience designing medical-accuracy or domain-specific evaluation suites.
- AI red-teaming or adversarial testing experience.
- Experience with production AI monitoring, drift detection, and continuous evaluation.
Similar jobs
- AT
Network Engineer
ATLAS TECH
Del Mar, CA🇺🇸Hybrid1 week agoAnsibleCDNPython+1Technology - SP
Lead Systems Engineer
NewSYSTEMS PLANNING AND ANALYSIS, INC.
Lorton, VA🇺🇸$170k - $200k/yrHybrid11 minutes agoTechnology - BA
Network Engineer
NewBOOZ, ALLEN & HAMILTON, INC.
Anne Arundel County, MD🇺🇸$99k - $225k/yrHybrid11 minutes agoEncryptionDNSZero Trust+1Technology - RT
Senior Software Engineer - Space and RF Sensors (Onsite) with Security Clearance
NewRTX
El Segundo, CA🇺🇸$95.5k - $181.7k/yrHybridYesterdayMachine LearningScrumAgile+7Technology - SN
Software Engineering Intern (Summer 2027) with Security Clearance
NewSierra Nevada Corporation
Plano, TX🇺🇸HybridYesterdayMATLABTechnology - RT
Sr. Principal System Engineer - Test Architect with Security Clearance
NewRTX
Tucson, AZ🇺🇸HybridYesterdayHTTPSTechnology