Evaluation Engineer (AI Models)
Quick Overview
Job Description
Job Description - Evaluation Engineer (AI Models)
Location: New York City, NY / Fort Mill, SC
Role Overview
We are seeking an experienced Evaluation Engineer - AI Models to join our growing AI and Digital Engineering team. The ideal candidate will have a strong background in Quality Engineering and hands-on experience evaluating AI/ML and Generative AI model performance across business and technical use cases.
This role requires a combination of analytical thinking, testing expertise, data-driven evaluation, and strong communication skills to collaborate effectively with engineering, product, and business stakeholders.
Key Responsibilities
- Design, develop, and execute evaluation strategies for AI/ML and Generative AI models.
- Validate model outputs for accuracy, relevance, consistency, hallucination detection, bias, safety, and performance.
- Create automated and manual evaluation frameworks for LLM-based applications and AI systems.
- Develop test cases, benchmarking approaches, and quality metrics for AI model validation.
- Work closely with Data Scientists, Product Managers, and Engineering teams to improve model quality and reliability.
- Analyze model behavior using structured and unstructured datasets.
- Perform regression testing and continuous validation for model updates and releases.
- Document evaluation findings, defects, risks, and recommendations clearly for technical and business audiences.
- Support UAT and production validation activities for AI-enabled products and platforms.
- Contribute to QA best practices, test automation strategies, and AI quality governance initiatives.
Required Qualifications
- 10+ years of experience in Quality Assurance / Quality Engineering / Software Testing.
- 2+ years of hands-on experience in AI model evaluation, Generative AI testing, or ML validation.
- Strong understanding of AI/ML concepts, LLM behavior, prompt evaluation, and model testing methodologies.
- Experience with API testing, test automation frameworks, and data validation techniques.
- Familiarity with evaluation metrics such as precision, recall, accuracy, grounding, relevance, and hallucination detection.
- Experience testing AI-powered applications, conversational AI, or GenAI platforms.
- Strong analytical and problem-solving skills.
- Excellent verbal and written communication skills.
- Ability to work independently in a remote and cross-functional environment.
Preferred Skills
- Experience with Python and AI/ML testing tools/frameworks.
- Exposure to prompt engineering and Retrieval-Augmented Generation (RAG) validation.
- Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Experience working in Agile/Scrum environments.
- Financial Services or Wealth Management domain experience is a plus.
Education
Bachelor s degree in Computer Science, Engineering, Information Systems, or related field preferred.
Skills
Similar jobs
Assembler I
Atlas Copco Group · Rock Hill, United States
11 minutes ago$1.5k/moAssociate Manager, Process Writing
Gainwell Technologies LLC · Austin, United States
13 minutes ago$104.1k - $113k/yrProcess Writer
BCforward · United States
14 minutes ago€30/hrRFP Team Member
Randstad Digital · Addison, United States
14 minutes ago$26 - $31/hrWorkshop Technician II
Atlas Copco Group · Hamburg, United States
14 minutes agoBusiness Execution Consultant 4
Motion Recruitment Partners, LLC · Charlotte, United States
14 minutes ago