Haystack
← Back to Jobs
Remote
Technology

Data Scientist

SAI Systems Intl., Inc.United States🇺🇸United StatesPosted 29 Jul 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Role: Data Scientist 
Location: Remote 
Job Summary
We are seeking a Senior AI/ML Engineer with strong expertise across Data Science, Generative AI/LLMs, AI Engineering, and Databricks. The ideal candidate will have hands-on experience selecting and developing AI/ML models, building RAG and vector-based solutions, deploying models into production, and leveraging Databricks, MLflow, and Feature Store for scalable ML workflows.
Key Responsibilities
Data Science & Model Development
  • Evaluate and select foundation/LLM models including Anthropic Claude (Opus), OpenAI GPT, Google Gemini, and other emerging models.
  • Develop, train, fine-tune, and optimize machine learning and GenAI models.
  • Perform hyperparameter tuning, model evaluation, benchmarking, and optimization.
  • Implement reinforcement learning/RLHF, human-in-the-loop (HITL), and feedback-driven model improvement.
  • Build programmatic model diagnostics, validation, controls, and quality monitoring.
  • Perform model drift analysis, data drift detection, performance monitoring, and remediation.
AI Engineering / GenAI
  • Design and implement end-to-end AI solutions from discovery through production.
  • Build RAG pipelines using embeddings, vector stores/vector databases, document processing, retrieval, reranking, and LLM generation.
  • Work with embeddings, vector search, semantic search, chunking, metadata filtering, and retrieval optimization.
  • Integrate LLMs with enterprise data, APIs, applications, and business workflows.
  • Automate model-to-data integration and model deployment across development, testing, and production environments.
  • Develop scalable AI services and APIs using Python and cloud-native technologies.
  • Implement LLM evaluation, observability, guardrails, and production monitoring.
Databricks / Cloud / MLOps
  • Develop and manage ML workflows using Databricks.
  • Strong hands-on experience with Databricks APIs, compute, MLflow, and Feature Store.
  • Build scalable data/ML pipelines using Apache Spark/PySpark.
  • Use MLflow for experiment tracking, model registry, model versioning, and deployment.
  • Implement feature engineering and manage reusable features through Feature Store capabilities.
  • Automate ML/AI deployment pipelines using CI/CD, MLOps, and cloud services.
  • Collaborate with data engineering, platform engineering, and application teams to productionize AI solutions.
Required Skills
  • Python, SQL, PySpark/Apache Spark
  • Generative AI, LLMs, RAG, Embeddings, Vector Databases
  • Hands-on experience with Claude/Anthropic, OpenAI GPT, Google Gemini
  • Model selection, fine-tuning, hyperparameter optimization
  • Reinforcement Learning / RLHF
  • HITL, model evaluation, diagnostics, drift analysis
  • Databricks, MLflow, Databricks APIs, Compute
  • Feature Store / Feature Engineering
  • MLOps, model deployment, CI/CD
  • Experience taking AI/ML solutions from POC/discovery → production
 

Skills

SQL
MLOps
MLflow
Machine Learning
Apache
Apache Spark
Databricks
GPT
Generative AI
LLM
Python

Similar jobs