Haystack
← Back to Jobs
Remote
Technology
BS

Principal Data Scientist- W2 | Remote

Blue Space TechnologiesUnited States🇺🇸United StatesPosted Sep 23, 2026

Quick Overview

Seniority
Leader
Work mode
Remote
Location
United States
Posted
19 hours ago
Machine LearningNLPDeep LearningGenerative AIKubernetesLLMPyTorchPythonReact

Job Description

Job Title: Principal Data Scientist Work Arrangement: Remote for the first 6 months, then onsite in Raleigh, NC
Client: LexisNexis

Position Overview

We are seeking a highly experienced Principal Data Scientist to lead the design, development, and deployment of advanced AI/ML and LLM-based systems. The ideal candidate will have strong experience in Data Science, Machine Learning Engineering, Generative AI, and production-grade LLM applications.

The candidate should be capable of working at both the research/architecture level and hands-on implementation level, translating advanced AI techniques into scalable enterprise solutions.

Required Qualifications

  • Master's degree or PhD preferred.
  • 10 12+ years of experience in Data Science / Machine Learning Engineering.
  • Deep hands-on experience with LLM-based systems and Generative AI.
  • Strong experience designing and implementing production-grade AI/ML solutions.
  • Experience with deep learning, NLP, transformers, and model fine-tuning.
  • Strong Python programming and ML engineering experience.
  • Experience building and deploying scalable ML/AI systems.

Key Technical Skills

Generative AI / LLM

  • Large Language Models (LLMs)
  • Generative AI
  • LLM application development
  • Prompt engineering
  • Model evaluation
  • Hallucination detection and mitigation
  • LLM observability and monitoring

RAG & Retrieval

  • Retrieval-Augmented Generation (RAG)
  • Embeddings
  • Vector databases
  • Semantic search
  • Hybrid retrieval
  • Reranking
  • Chunking strategies
  • Metadata filtering
  • Retrieval optimization

Machine Learning / Deep Learning

  • PyTorch
  • HuggingFace
  • Transformers
  • NLP
  • Deep learning
  • Model fine-tuning
  • Classification
  • Information extraction
  • Summarization
  • Question answering

AI Agent Architecture

  • Multi-agent architectures
  • Planner-executor patterns
  • Tool-use agents
  • ReAct-style reasoning
  • Agent orchestration
  • Tool calling
  • Workflow routing
  • Memory and guardrails

Frameworks & Platforms

  • LangChain
  • LlamaIndex
  • Kubernetes
  • Containerized model serving
  • Production ML APIs
  • Monitoring and observability
  • Cloud/production deployment

Data & Vector Technologies

Experience with technologies such as:

  • Pinecone
  • Weaviate
  • FAISS
  • Chroma
  • pgvector
  • Embedding indexes
  • Vector retrieval infrastructure

Responsibilities

  • Lead the architecture and development of enterprise-scale AI/ML and LLM solutions.
  • Design and implement production-grade RAG pipelines.
  • Develop and optimize multi-agent AI architectures.
  • Build LLM applications using frameworks such as LangChain and LlamaIndex.
  • Fine-tune and evaluate transformer-based models.
  • Develop evaluation frameworks for model quality, retrieval accuracy, groundedness, latency, cost, and reliability.
  • Establish methods for detecting hallucinations and other LLM failure modes.
  • Design scalable ML serving and deployment architectures.
  • Work with Kubernetes and containerized inference environments.
  • Implement monitoring, observability, alerting, and model-performance tracking.
  • Collaborate with Data Scientists, ML Engineers, Product teams, and other technical stakeholders.
  • Provide technical leadership, architecture guidance, mentoring, and code reviews.
  • Translate business requirements into scalable AI/ML solutions.
  • Drive AI strategy, technical roadmaps, experimentation, and continuous improvement.

Preferred Experience

  • Experience leading principal-level AI/ML architecture.
  • Experience deploying LLM systems into production at enterprise scale.
  • Experience with large-scale data and knowledge-retrieval systems.
  • Experience creating measurable AI evaluation frameworks.
  • Experience with model monitoring, observability, and production reliability.
  • Strong communication and stakeholder-management skills.
  • Experience mentoring Data Scientists and ML Engineers.

Work Arrangement

  • First 6 months: Remote
  • After 6 months: Onsite in Raleigh, NC
  • Candidates must be willing to relocate to Raleigh, NC after the initial remote period.

Similar jobs