Haystack
← Back to Jobs
Technology
UI

AI Data Engineer

United IT SolutionsUnited States🇺🇸United StatesPosted 31 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
23 hours ago
DockerETLMachine LearningDatabricksGitHub ActionsJenkinsKubernetesLLM

Job Description

Job Overview

We are seeking an AI Data Engineer who bridges the gap between Data Science and Data Engineering. The ideal candidate understands how to take Machine Learning and LLM prototypes/concepts and convert them into robust, scalable, production-grade data pipelines. While deep Machine Learning model training is not the primary focus, a foundational understanding of ML, LLM architectures, and RAG frameworks is required to design and optimize end-to-end AI pipelines.

Key Responsibilities

·         Pipeline Engineering: Design, build, and maintain production-level data pipelines to deploy and operationalize ML and LLM workflows.

·         LLM & RAG Integration: Implement Retrieval-Augmented Generation (RAG) frameworks using libraries like LangChain or LlamaIndex to query structured and unstructured data sources.

·         API & System Integration: Integrate LLM APIs (e.g., OpenAI, Anthropic, or open-source models) into data processing workflows.

·         Performance Optimization: Optimize distributed workloads, data processing engines, and pipeline latency for real-time and batch execution.

·         CI/CD & DevOps: Build and maintain CI/CD pipelines to deploy data and AI workflows using relevant SDKs and automation tools.

·         Data Quality & Validation: Implement strict schema validation rules and data quality checks to ensure reliable pipeline execution.

·         AI Evaluation & Quality Control: Monitor and measure output quality using key metrics such as retrieval quality, answer correctness, and faithfulness.

 

Required Qualifications

·         Experience: Proven experience as a Data Engineer building production-grade ETL/ELT data pipelines.

·         LLM / AI Concepts: Minimum working knowledge of ML concepts, LLM architectures, vector databases, and RAG frameworks (e.g., LangChain, LlamaIndex).

·         API Integration: Hands-on experience integrating third-party or self-hosted LLM APIs into data pipelines.

·         Distributed Computing: Experience optimizing distributed data processing workloads (e.g., PySpark, Spark, Databricks, Ray, or Cloud-native processing services).

·         CI/CD & Automation: Solid understanding of CI/CD pipeline implementation, deployment SDKs, and containerization (e.g., Docker, Kubernetes, GitHub Actions, Jenkins).

·         Schema & Data Quality: Expertise in enforcing schema validation rules, data contracts, and pipeline performance optimization.

·         Evaluation Metrics: Familiarity with AI/RAG evaluation metrics (e.g., retrieval precision, answer correctness, context relevance, faithfulness).

Similar jobs