Haystack
← Back to Jobs
Technology
CL

Senior Data Scientist – Generative AI / LLM Evaluation

ClifyXWestwood, MA🇺🇸United StatesPosted Sep 16, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Westwood, MA, United States
Posted
19 hours ago
AWSMLflowMachine LearningNumPyScikit-learnDatabricksGenerative AILLMPandasPython

Job Description

Job Title: Senior Data Scientist – Generative AI / LLM Evaluation

Work Location: Johnston, RI or Westwood, MA (Hybrid)/Any Infosys Hub Office

Contract duration: Long term contract  

Visa: / Only

Domain (Industry)         

Cards, Banking, FSI

 

Position Summary

Job Details:

We are seeking an experienced Data Scientist to support the development and evaluation of AI-powered fraud self-service voice agents and conversational AI systems. The primary responsibility is not model deployment or engineering implementation, but designing evaluation frameworks, measuring system performance, identifying failure patterns, conducting root-cause analysis, and optimizing model behavior through data-driven experimentation.

Key Responsibilities

  • Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems.
  • Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes.
  • Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation.
  • Perform prompt tuning and experimentation to improve model accuracy and reliability.
  • Conduct root-cause analysis of conversational failures and recommend remediation strategies.
  • Measure performance across different model configurations, prompts, and guardrail implementations.
  • Partner with AI Engineering and Product teams to validate solutions before production deployment.
  • Build dashboards and reports that communicate model effectiveness and operational impact.
  • Support fraud-related customer service use cases, including intent detection and multi-turn conversation flows.

Success Criteria

  • Develop reliable evaluation methodologies for conversational AI systems.
  • Quantify the effectiveness of fraud self-service voice agents.
  • Optimize prompts, retrieval strategies, and guardrails using empirical evidence.
  • Deliver actionable insights that improve customer experience and model performance.
  • Establish measurable KPIs for intent detection and multi-turn conversation success.

 

Mandatory skills:

  • Strong background in data science, machine learning, generative AI, or a related quantitative field.
  • Hands-on experience evaluating LLM, RAG, agentic AI, or conversational AI solutions.
  • Deep understanding of model evaluation techniques and metrics, including:
  • Precision@K
  • Recall@K
  • Mean Reciprocal Rank (MRR)
  • F1 Score
  • Retrieval and generation quality assessment
  • Experience performing experimentation, statistical analysis, and performance benchmarking.
  • Strong Python programming skills.
  • Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy, and related tools.
  • Ability to communicate technical findings succinctly to highly technical stakeholders.

 

Important Note

This role is primarily a Data Science and AI Evaluation position, not an AI Engineering or deployment-focused role. The emphasis is on measuring, analyzing, validating, and improving AI system performance rather than building production deployment pipelines.

 

Desired skills:

  • Experience with:
  • Generative AI and LLM ecosystems
  • Multi-agent systems
  • RAG/Agentic RAG architectures
  • Amazon Bedrock
  • AWS SageMaker
  • Databricks
  • MLflow
  • LangSmith
  • Weights & Biases
  • Knowledge of conversational AI, IVR systems, digital assistants, and voice agents.
  • Experience in financial services, fraud detection, or customer service automation.

Similar jobs