Haystack
← Back to Jobs
Technology

AI Lead Data Engineer (No GC''s)

Millennium Global TechnologiesDallas, TX🇺🇸United StatesPosted 21 Jul 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

CW - Lead Data Engineer (T4I Data Team) 

Location: Mountain View,CA/ Dallas,TX 

Onsite 3 days office

 

 8-11 year experience

This is a mandatory requirement. It is not sufficient for these skills to simply appear on the resume they must be backed by hands-on, project-level experience.

The candidate must have proven hands-on experience with AI toolssemantic layer development, and AWSCandidates whose experience is primarily with Azure will not be considered, 

 

 

About the Role T4I''''s data team builds and operates the data marts and pipelines that power HR/Workforce reporting and analytics 

We''''re looking for a hands-on Data Engineer who can operate independently. 

High-visibility, high-ownership role at the center of T4I''''s HR data platform direct exposure to AI initiatives, cross-functional stakeholders, and the chance to meaningfully reduce day-to-day operational load for the team.

 

What You''''ll Do 

• Build, maintain, and troubleshoot Spark/EMR ETL pipelines feeding HR and Workforce data marts. 

• Monitor and remediate Data Asset Score issues (data quality rules, governance/PII-minimization actions) to keep HR data assets compliant ahead of deadlines.

• Be the daily coordination point between US HR business stakeholders, the IDC (India) engineering team, and platform/infra teams - translating requirements, unblocking IDC, and reporting status. 

• Triage and resolve Jira tickets/bugs raised against HR data mart pipelines; write clear runbooks and pipeline documentation. 

• Build semantic layers and Retrieval-Augmented Generation (RAG) pipelines. 

• Integrate REST and/or GraphQL APIs into data workflows. 

• Operate with minimal oversight: escalating only true blockers and proactively flag risks before they become incidents. What You Bring

• 

 

8+ years in Data Engineering, with strong hands-on Spark (PySpark/Scala), SQL, and Python.

• Experience building and operating ETL/ELT pipelines on cloud platforms (AWS EMR/S3, Databricks, or equivalent); workflow orchestration (Airflow or similar).

• Comfortable owning pipeline operations end-to-end: reading dependency graphs, diagnosing failures, working with on-call /PagerDuty/Splunk, and driving fixes with multiple teams.

• Working knowledge of data warehousing/data mart design, data governance, and PII handling practices. 

• AI-native mindset: familiarity with LLM capabilities, evaluation frameworks, and the creative application of AI principles to engineering challenges. • Proficiency with GenAI productivity tools (e.g., Claude, Cursor, Codex) to enhance engineering workflows. 

• Strong communicator; comfortable directing an offshore IDC team with minimal handholding.

Skills

SQL
Scala
AWS
ETL
Splunk
Airflow
Azure
Databricks
GraphQL
Jira
LLM
PagerDuty
Python
REST

Similar jobs