Quick Overview
Job Description
Job Overview
We are seeking an AI Data Engineer who bridges the gap between Data Science and Data Engineering. The ideal candidate understands how to take Machine Learning and LLM prototypes/concepts and convert them into robust, scalable, production-grade data pipelines. While deep Machine Learning model training is not the primary focus, a foundational understanding of ML, LLM architectures, and RAG frameworks is required to design and optimize end-to-end AI pipelines.
Key Responsibilities
· Pipeline Engineering: Design, build, and maintain production-level data pipelines to deploy and operationalize ML and LLM workflows.
· LLM & RAG Integration: Implement Retrieval-Augmented Generation (RAG) frameworks using libraries like LangChain or LlamaIndex to query structured and unstructured data sources.
· API & System Integration: Integrate LLM APIs (e.g., OpenAI, Anthropic, or open-source models) into data processing workflows.
· Performance Optimization: Optimize distributed workloads, data processing engines, and pipeline latency for real-time and batch execution.
· CI/CD & DevOps: Build and maintain CI/CD pipelines to deploy data and AI workflows using relevant SDKs and automation tools.
· Data Quality & Validation: Implement strict schema validation rules and data quality checks to ensure reliable pipeline execution.
· AI Evaluation & Quality Control: Monitor and measure output quality using key metrics such as retrieval quality, answer correctness, and faithfulness.
Required Qualifications
· Experience: Proven experience as a Data Engineer building production-grade ETL/ELT data pipelines.
· LLM / AI Concepts: Minimum working knowledge of ML concepts, LLM architectures, vector databases, and RAG frameworks (e.g., LangChain, LlamaIndex).
· API Integration: Hands-on experience integrating third-party or self-hosted LLM APIs into data pipelines.
· Distributed Computing: Experience optimizing distributed data processing workloads (e.g., PySpark, Spark, Databricks, Ray, or Cloud-native processing services).
· CI/CD & Automation: Solid understanding of CI/CD pipeline implementation, deployment SDKs, and containerization (e.g., Docker, Kubernetes, GitHub Actions, Jenkins).
· Schema & Data Quality: Expertise in enforcing schema validation rules, data contracts, and pipeline performance optimization.
· Evaluation Metrics: Familiarity with AI/RAG evaluation metrics (e.g., retrieval precision, answer correctness, context relevance, faithfulness).
Similar jobs
- BT
Master Data Engineer | Configuration Management
BETA Technologies
South Burlington🇺🇸$80k - $110k/yr3 days agoCADCATIAERP+3Technology - AS
Data Engineer 3
NewApex Systems
Atlanta, GA🇺🇸$55 - $60/hrOn-site23 hours agoSQLSQL ServerETL+5Technology - ME
Senior Data Engineer (Agentic AI / RAG Platform) - Onsite - FORD- Locals Only
NewMeganSoft
Dearborn, MI🇺🇸On-site23 hours agoSQLAzurePythonTechnology - LM
Azure Data Engineer with AI Experience : Remote
NewLightning Minds Inc.
United States🇺🇸Remote23 hours agoSQLAzurePostgreSQL+2Technology - BA
Senior Help Desk Data Engineer
NewBooz Allen Hamilton
Winchester, VA🇺🇸$61.9k - $141k/yrOn-site23 hours agoMySQLOracleRust+16Technology - RI
Data Engineer
NewRITWIK Infotech Inc
Dallas, TX🇺🇸Hybrid23 hours agoSQLLookerData Pipeline+4Technology