Haystack
← Back to Jobs
Remote
Technology
TT

AI Data Engineer

THE TILTED CIRCLE LLCUnited States🇺🇸United StatesPosted 3 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
20 hours ago
Neo4jSQLAWSMLflowMachine LearningScikit-learnApacheApache SparkAzureDatabricksGenerative AIGoogle CloudPyTorchPythonTensorFlow

Job Description

Position: AI Data Engineer

Location: Remote/NJ, NY (Hybrid)

Duration: Full Time

 

Hybrid in Princeton, NJ/NYC, NY, or Remote

 

Hybrid is Preferred or else look for candidate in EAST who can work Remotely and can come onsite once a month with their own expenses

 

Please submit along with availability for next 3 days multiple slots after 3 PM EST

 

Role Overview

The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.

 

Key Responsibilities

  • Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.
  • Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.
  • Develop MCP servers and enable AI data distribution via MCP.
  • Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.
  • Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.
  • Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.
  • Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.
  • Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
  • Develop data quality frameworks specific to AI training datasets and validation data.
  • Collaborate with data scientists to optimize data access patterns and feature store implementations.
  • Implement security and compliance controls for sensitive AI training data and model artifacts.
  • Create comprehensive documentation for AI data architectures, schemas, and integration patterns.

 

Required Skills and Qualifications

  • Bachelor's or master’s degree in computer science, Data Science, Machine Learning, or related field.
  • 5-7 years of hands-on experience in data engineering, with at least 2 years focused on AI/ML workloads.
  • Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.
  • Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.
  • Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.
  • Experience with feature engineering, data preprocessing techniques, and ML data pipelines.
  • Knowledge of vector databases, embeddings, and similarity search for AI applications.
  • Proficiency in SQL for structured and unstructured data management.
  • Understanding of data governance, model governance, and AI ethics principles.
  • Strong analytical and problem-solving capabilities with attention to data quality.
  • Excellent collaboration skills for working with data scientists, ML engineers, and architects.

 

Preferred/Nice-to-Have Skills

  • Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine-tuning.
  • Knowledge of LangChain, HuggingFace, or other GenAI frameworks.
  • Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.
  • Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.
  • Understanding of AI model explainability and interpretability techniques.
  • Experience with A/B testing frameworks for ML model evaluation.
  • Certification in Databricks, AWS, Azure, or Google Cloud Platform AI/ML services.
  • Publications or contributions to open-source ML projects

Similar jobs