Haystack
← Back to Jobs
Technology
TR

Principal Machine Learning Engineer Agentic AI Platforms

Tech RakersRaleigh, NC🇺🇸United StatesPosted Sep 29, 2026

Quick Overview

Seniority
Leader
Work mode
On Site
Location
Raleigh, NC, United States
Posted
18 hours ago
DockerFastAPISQLAWSMLOpsMachine LearningAzureGoogle CloudKubernetesLLMPythonREST

Job Description

Principal Machine Learning Engineer Agentic AI Platforms

2 days onsite in Raleigh, NC

Long Term

Position Summary

We are seeking a Principal Machine Learning Engineer with 10+ years of experience building large-scale distributed machine learning systems and modern AI platforms. This role will lead the architecture and implementation of Agentic AI ecosystems, LLM infrastructure, RAG platforms, model-serving systems, and AI engineering frameworks supporting enterprise-wide AI adoption.

This is a deeply hands-on role requiring expertise in platform architecture, ML systems design, distributed computing, MLOps, LLMOps, and production AI deployment.

Key Responsibilities

Agentic AI Platform Engineering

  • Design and implement enterprise-grade agent frameworks supporting:
    • Multi-agent collaboration
    • Planner-executor architectures
    • Tool-use workflows
    • Memory systems
    • Human-review workflows
  • Build orchestration services and runtime environments for AI agents.
  • Implement MCP and agent interoperability standards.
  • Design state-sharing and context management systems.

LLM Platform Development

  • Build scalable LLM platforms supporting:
    • OpenAI
    • Claude
    • Gemini
    • Llama
    • Mistral
    • Fine-tuned models
  • Implement model routing and workload balancing across models.
  • Design low-latency AI serving architectures.

RAG & Retrieval Engineering

  • Build retrieval systems using:
    • Hybrid search
    • Dense retrieval
    • BM25
    • Re-ranking pipelines
    • Graph-based retrieval
  • Optimize vector stores and embedding pipelines.
  • Design retrieval infrastructure capable of supporting billions of documents.

MLOps & LLMOps

  • Build automated pipelines for:
    • Training
    • Evaluation
    • Deployment
    • Monitoring
    • Retraining
  • Implement CI/CD pipelines for AI systems.
  • Build observability frameworks including:
    • OpenTelemetry
    • LangSmith
    • LangFuse
    • Custom tracing

Reliability & Evaluation

  • Implement release gates and quality controls.
  • Develop evaluation pipelines for:
    • Hallucination detection
    • Retrieval quality
    • Agent performance
    • Cost optimization
    • Latency management
  • Design monitoring frameworks for drift, degradation, and reliability.

Technical Leadership

  • Lead architecture reviews for AI systems.
  • Establish engineering standards and best practices.
  • Mentor engineering teams.
  • Partner with Data Scientists, Product Managers, and Business Stakeholders.

Required Qualifications

  • Bachelor's, Master's, or PhD in Computer Science, Engineering, Data Science, or related field.
  • 10+ years of software engineering or machine learning engineering experience.
  • 5+ years building production ML systems.
  • 3+ years building production GenAI platforms.
  • Strong expertise in:
    • Python
    • SQL
    • Spark
    • Distributed computing
    • Kubernetes
    • Docker

Experience with:

  • LangGraph
  • LangChain
  • LlamaIndex
  • Vector databases
  • FastAPI
  • REST APIs
  • Cloud platforms (AWS, Azure, Google Cloud Platform)

Preferred Qualifications

  • Experience building multi-agent platforms at enterprise scale.
  • Experience implementing MCP and A2A architectures.
  • Knowledge graph and GraphRAG experience.
  • Fine-tuning experience (LoRA, QLoRA, PEFT).
  • Experience deploying AI systems in healthcare, finance, legal, insurance, or cybersecurity domains.

Sign Ideal Candidate Profile

The strongest candidates for either role should be able to confidently discuss:

  • Agent architecture tradeoffs
  • Retrieval evaluation methodologies
  • Hallucination measurement
  • LLM evaluation frameworks
  • Fine-tuning vs RAG decisions
  • Production AI failures and learnings
  • Cost optimization strategies
  • MCP/A2A interoperability
  • AI governance and model risk management
  • Scaling AI systems from prototype to enterprise production
  • ificant experience operating AI systems supporting thousands to millions of users.

Similar jobs