Why This Role Stands Out
This hybrid role offers a fantastic opportunity to shape the future of AI data infrastructure, working with cutting-edge technologies at a reputable company. You'll thrive here if you're a skilled data engineer with a passion for AI and a desire for impactful work in a collaborative environment. Apply now to advance your career and contribute to innovative AI solutions!
Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
New York, NY, United States
Posted
19 hours ago
Neo4jSQLAWSMLflowMachine LearningScikit-learnApacheApache SparkAzureDatabricksGenerative AIGoogle CloudPyTorchPythonTensorFlow
Job Description
Senior AI Data Engineer
Location- Princeton, NJ & NYC, NY (Hybrid)
AI Data Engineer
Role Overview
The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.
Key Responsibilities
Location- Princeton, NJ & NYC, NY (Hybrid)
AI Data Engineer
Role Overview
The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.
Key Responsibilities
- Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.
- Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.
- Develop MCP servers and enable AI data distribution via MCP.
- Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.
- Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.
- Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.
- Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.
- Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.
- Develop data quality frameworks specific to AI training datasets and validation data.
- Collaborate with data scientists to optimize data access patterns and feature store implementations.
- Implement security and compliance controls for sensitive AI training data and model artifacts.
- Create comprehensive documentation for AI data architectures, schemas, and integration patterns.
- Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.
- 5-7 years of hands-on experience in data engineering, with at least 2 years focused on AI/ML workloads.
- Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.
- Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.
- Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.
- Experience with feature engineering, data preprocessing techniques, and ML data pipelines.
- Knowledge of vector databases, embeddings, and similarity search for AI applications.
- Proficiency in SQL for structured and unstructured data management.
- Understanding of data governance, model governance, and AI ethics principles.
- Strong analytical and problem-solving capabilities with attention to data quality.
- Excellent collaboration skills for working with data scientists, ML engineers, and architects.
- Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine-tuning.
- Knowledge of LangChain, HuggingFace, or other GenAI frameworks.
- Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.
- Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.
- Understanding of AI model explainability and interpretability techniques.
- Experience with A/B testing frameworks for ML model evaluation.
- Certification in Databricks, AWS, Azure, or Google Cloud Platform AI/ML services.
- Publications or contributions to open-source ML projects
Similar jobs
- EX
Sr. Data Engineer- FULLTIME Role / NYC NY.(Hybrid 34 days onsite)
NewExatech Inc
New York, NY🇺🇸On-site19 hours agoSQLETLScrum+5Technology - MM
Senior Data Quality Engineer
NewMitchell Martin, Inc.
Chicago, IL🇺🇸$59 - $84/hrOn-site19 hours agoEngineering - AA
DATA ENGINEER
NewAaraTechnologies Inc
Jersey City, NJ🇺🇸Hybrid19 hours agoSQLAWSETL+2Technology - TT
TMAS3 Tenants Mission Data Engineer with Security Clearance
NewTorch Technologies Inc.
Eglin AFB, FL🇺🇸Hybrid19 hours agoTechnology - DW
AI/ML Data Engineer at Princeton, NJ & NYC, NY (Hybrid)
NewData Wave Technologies Inc
Princeton, NJ🇺🇸Hybrid19 hours agoSQLMachine LearningScikit-learn+6Technology - ST
Senior Data Engineer
NewSSV Technologies Inc
United States🇺🇸Remote19 hours agoSQLAWSETL+11Technology