Haystack
← Back to Jobs
Technology

AI Data Engineer

QUANTUM TECHNOLOGIES LLCTallahassee, FL🇺🇸United StatesPosted 23 Jul 2026

Why This Role Stands Out

This AI Data Engineer role offers a fantastic opportunity to shape cutting-edge AI and LLM integrations within a reputable company, with highly competitive hourly compensation. You'll thrive here if you're passionate about building robust data pipelines and eager to develop advanced skills in areas like RAG and Generative AI. Apply today to contribute your expertise to innovative projects!

Quick Overview

Salary
€90/hr
Work Type
On Site
Level
Mid Senior

Job Description

AI Data Engineer

Location: Tallahassee, FL, USA

Duration: 12 Months + Extension

Bill Rate: $90/hr on C2C

Job Type: C2C/1099 Contract

Client: To Be Discussed Later

Work Authorization: US-Citizen, H-1B, OPT-EAD, GC-EAD

Job Description:

  • Design, develop, and optimize scalable batch and real-time data pipelines using Apache Spark (PySpark/Spark SQL).
  • Write complex, high-performance SQL queries for data extraction, transformation, analytics, and reporting.
  • Build and maintain ETL/ELT pipelines to ingest, cleanse, transform, and integrate structured and unstructured data.
  • Prepare, curate, and validate datasets for Machine Learning and Generative AI applications.
  • Develop and optimize RAG (Retrieval-Augmented Generation) data pipelines using vector databases and document processing frameworks.
  • Integrate enterprise data with Large Language Models (LLMs) such as OpenAI GPT, Azure OpenAI, Claude, or Gemini.
  • Implement AI-powered data quality validation, anomaly detection, and automated monitoring solutions.
  • Perform data validation, testing, and quality assurance to ensure data accuracy, completeness, and consistency.
  • Optimize Spark jobs, SQL queries, and distributed processing for maximum performance and scalability.
  • Collaborate with Data Scientists, AI Engineers, Business Analysts, and Application Developers to support AI initiatives.
  • Monitor ETL workflows and AI data pipelines to ensure reliable and timely data delivery.
  • Maintain technical documentation for data architecture, AI pipelines, metadata, and data models.
  • Follow best practices for data governance, security, compliance, and AI model lifecycle management.
    Preferred Qualifications:
  • 5+ years of experience as a Data Engineer.
  • Strong hands-on experience with SQL and query optimization.
  • Extensive experience with Apache Spark (PySpark/Spark SQL).
  • Strong Python programming experience.
  • Experience building and maintaining scalable ETL/ELT pipelines.
  • Strong understanding of relational databases and data modeling.
  • Experience with Generative AI, LLMs, and AI-powered data engineering workflows.
  • Knowledge of RAG architecture, embeddings, vector databases (Pinecone, ChromaDB, FAISS, or Weaviate), and semantic search.
  • Experience with AI frameworks such as LangChain or LlamaIndex.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Experience with data testing, validation, and quality assurance.
  • Strong analytical, troubleshooting, and communication skills.
  • Ability to work effectively in an onsite, collaborative environment.
  • Experience with Databricks.
  • Experience with Delta Lake, Apache Iceberg, or Apache Hudi.
  • Knowledge of Kafka, event-driven architectures, or streaming data pipelines.
  • Experience with Airflow or other workflow orchestration tools.
  • Experience with Docker and Kubernetes.
  • Familiarity with MLOps tools such as MLflow.
  • Experience implementing enterprise AI governance and responsible AI practices.

Skills

Docker
SQL
AWS
ETL
MLOps
MLflow
Machine Learning
Airflow
Apache
Apache Spark
Azure
Databricks
GPT
Generative AI
Google Cloud
Kafka
Kubernetes
Python

Similar jobs