Quick Overview
Job Description
Role: Lead Data Engineer (Hands-On)
Location: Cary, NC (On-site / Hybrid)
Experience: 12-18 years
Employment: Full-Time
Salary: $140K - $145K per annum plus benefits
ABOUT THE ENGAGEMENT
A centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks. The platform ingests
150+ inbound data feeds, distributes to 35+ downstream systems, and is organized as a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business user access from day one.
This is a senior hands-on leadership role. The candidate will own the end-to-end technical design of the data and AI layers, build reference implementations for the engineering team, and ship production-grade
Python, Scala, and PySpark code every week. Candidates who have not written or reviewed production code in the past year are not a fit. WHAT THE ROLE OWNS
- Data platform architecture and engineering: lakehouse architecture
- Metadata-driven, parameterized ingestion frameworks for batch files, database extracts, CDC feeds and streaming (Azure Event Hubs / Kafka,
- Canonical PySpark and Scala Spark jobs, coding and testing standards, PR reviews, production incident debugging, Spark cluster tuning and cost guardrails.
- CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset
Analytics.
- AI-augmented ingestion and canonical mapping: auto-generated bridge documents, DML, canonical table definitions; AI-assisted source-to-canonical mapping with human review gate.
- AI-driven data quality, anomaly detection (data drift, schema drift, volume shifts, reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data.
- Semantic layer and knowledge graph, plus a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security.
- Governance and leadership: Unity Catalog (lineage, access control,
MUST-HAVE SKILLS & EXPERIENCE
- Expert-level Python, Scala and PySpark: production-ready, modular, well-tested solutions; Spark workload troubleshooting; optimizing large-scale batch and streaming pipelines using Delta Lake.
- Strong SQL and data modelling (dimensional and normalised), schema design, data contracts.
- Databricks expertise: Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.
- Azure data stack: ADLS Gen2 (zone design, ACLs, lifecycle), Azure
Hubs.
- 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, prompt engineering.
- Evaluation discipline: golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback.
- Hands-on with LangChain, LlamaIndex or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
- Metadata-driven frameworks: schema inference, data profiling, lineage, catalogs.
- 12-18 years of total experience in data engineering / data platform delivery.
- Proven enterprise-scale delivery of a medallion / lakehouse architecture.
- Azure security and governance: Entra ID, managed identities, RBAC,
- CI/CD and IaC: Azure DevOps, Terraform, Databricks Asset Bundles, automated testing of data pipelines.
- Clear technical writing and ability to present and defend designs to engineers and non-technical stakeholders.
- Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modelling over a lakehouse).
- Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale.
- ML-based anomaly detection on time-series or transactional financial data.
- Financial services or insurance domain (finance close, GL, subledger, reconciliation, actuarial data).
- LLMOps / MLOps: model and prompt versioning, cost governance, observability.
- Databricks Data Engineer Professional, Azure DP-203 / DP-700, or
- dbt, Great Expectations or similar; Workday, Prism or Accounting
Regards,
Martin
Similar jobs
- BA
Palantir Data Engineer with Security Clearance
NewBooz Allen Hamilton
Aberdeen Proving Ground, MD🇺🇸$77.6k - $176k/yrOn-site19 hours agoMachine LearningAgileGit+3Technology - NI
Database Engineer with Security Clearance
NewNiyamIT Inc
Ashburn, VA🇺🇸Remote19 hours agoMongoDBOracleSQL+10Technology - CA
Data Migration Engineer
Capgemini America, Inc.
United States🇺🇸$80.4k - $106.0k/yrHybrid1 week agoEngineering - DU
Senior Data Engineer (FedD238)
NewDefense Unicorns
Springfield, Massachusetts🇺🇸$148.8k - $201.3k/yrOn-site5 hours agoGCPOracleSQL+11Technology - AC
Sr. AI & Data Engineer - Databricks
NewAccenture
Dallas, Texas🇺🇸$70.3k - $205.8k/yrHybrid5 hours agoSQLMLOpsApache+7Technology - DU
FDE Data Engineer- Space (FedD141/FedD147)
NewDefense Unicorns
United States🇺🇸$123.3k - $166.8k/yrRemote5 hours agoSOAPSQLFlink+10Technology - AC
Sr. AI & Data Engineer - Snowflake
NewAccenture
Dallas, Texas🇺🇸$70.3k - $205.8k/yrHybrid5 hours agoGCPSQLAWS+13Technology - AS
Senior Data Engineer
NewApex Systems
Lemont, IL🇺🇸Hybrid19 hours agoSQLETLStakeholder ManagementTechnology - AS
Data Engineer 3
NewApex Systems
Boise, ID🇺🇸Hybrid19 hours agoSQLSQL ServerETL+4Technology - BA
Data Engineer
NewBooz Allen Hamilton
Fort Belvoir, VA🇺🇸$77.6k - $176k/yrOn-site19 hours agoPowerShellPythonZero TrustTechnology - BA
Data Engineer, Senior with Security Clearance
NewBooz Allen Hamilton
Arlington, VA🇺🇸$77.6k - $176k/yrOn-siteYesterdaySQLAWSETL+2Technology - BA
Data Engineer, Senior with Security Clearance
NewBooz Allen Hamilton
Charlottesville, VA🇺🇸$77.5k - $176k/yrOn-siteYesterdayDockerMongoDBMySQL+24Technology