Haystack
← Back to Jobs
Technology

W2 - Senior Data Scientist / ML Engineer (Generative AI)

ProhiresBoston, MA🇺🇸United StatesPosted 11 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Position: Senior Data Scientist / ML Engineer (Generative AI)

Location: Boston, MA (Onsite)

Duration: Contract & Full time

  

Primary Objective

 

We are seeking a Senior Data Scientist / ML Engineer specializing in Generative AI to design, evaluate, optimize, and productionize AI/ML solutions for enterprise applications—including RAG systems, AI agents, intelligent automation, and model evaluation platforms.

 

The role focuses on improving AI accuracy and retrieval quality, reducing hallucinations, benchmarking LLMs, and building reliable solutions for enterprise-scale deployments.

 

Success looks like: measurable gains in retrieval/answer quality, robust evaluation frameworks in production, and clear collaboration with AI engineering to ship governed, reliable GenAI systems.

 

Key Responsibilities

 

Primary

Design and develop machine learning and Generative AI solutions.

Build and optimize RAG pipelines, retrieval strategies, embeddings, and semantic search.

Evaluate and benchmark LLMs for accuracy, performance, and reliability.

Develop AI evaluation frameworks for hallucination detection, accuracy measurement, bias/toxicity detection, and ground-truth validation.

Optimize prompts, models, and retrieval workflows.

Collaborate with AI engineering teams to deploy models into production.

Also expected

Create training, validation, and testing datasets.

Perform model benchmarking, A/B testing, and performance analysis.

Fine-tune foundation models when required.

Implement model monitoring, observability, and ongoing evaluation processes.

 

Must-Have Experience & Skills

5–10 years of experience in data science, machine learning, or related applied ML roles.

Strong Python programming skills, with Pandas, NumPy, and Scikit-learn.

Strong foundation in supervised/unsupervised learning, statistical modeling, feature engineering, and model evaluation techniques.

Hands-on Generative AI experience with LLM evaluation, prompt engineering, RAG architectures, embedding models, fine-tuning approaches, and agent evaluation frameworks.

Experience with PyTorch and/or TensorFlow.

Exposure to OpenAI models, Claude, Gemini, and/or open-source LLMs.

Experience with vector databases, semantic search, and retrieval optimization.

Experience delivering or supporting production AI/ML solutions in enterprise environments.

Experience working with distributed onshore/offshore teams.

 

Preferred Skills

Databricks, MLflow, and Spark.

GraphRAG and Knowledge Graphs; exposure to Neo4j.

Responsible AI, Explainable AI, and AI governance.

Banking or Financial Services domain experience.

Familiarity with Azure AI Foundry, AWS Bedrock, Kubernetes, and AI observability platforms.

Soft Skills

Clear communication with engineering and business stakeholders.

Ability to translate evaluation results into actionable model/product decisions.

Comfortable owning quality metrics and trade-offs (accuracy, latency, cost, risk) in a delivery setting.

 

Skills

Neo4j
AWS
MLflow
Machine Learning
NumPy
Scikit-learn
Azure
Databricks
Generative AI
Kubernetes
LLM
Pandas
PyTorch
Python
TensorFlow

Similar jobs