Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Cary, NC, United States
Posted
18 hours ago
Neo4jSQLScalaMLOpsAzureDatabricksGPTKafkaLLMPythonTerraformUnityVaultdbt
Job Description
Onsite/Hybrid Position
Fulltime Position
Experience: 12-18 years
ABOUT THE ENGAGEMENT
A centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks. The platform ingests 150+ inbound data feeds, distributes to 35+ downstream systems, and is organized as a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business user access from day one.
This is a senior hands-on leadership role. The candidate will own the end-to-end technical design of the data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code every week. Candidates who have not written or reviewed production code in the past year are not a fit.
WHAT THE ROLE OWNS
A centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks. The platform ingests 150+ inbound data feeds, distributes to 35+ downstream systems, and is organized as a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business user access from day one.
This is a senior hands-on leadership role. The candidate will own the end-to-end technical design of the data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code every week. Candidates who have not written or reviewed production code in the past year are not a fit.
WHAT THE ROLE OWNS
- - Data platform architecture and engineering: lakehouse architecture (Bronze / Silver / Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, retention).
- - Metadata-driven, parameterized ingestion frameworks for batch files, database extracts, CDC feeds and streaming (Azure Event Hubs / Kafka, Spark Structured Streaming).
- - Canonical PySpark and Scala Spark jobs, coding and testing standards, PR reviews, production incident debugging, Spark cluster tuning and cost guardrails.
- - CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset Bundles and Terraform; observability with Azure Monitor and Log Analytics.
- - AI-augmented ingestion and canonical mapping: auto-generated bridge documents, DML, canonical table definitions; AI-assisted source-to-canonical mapping with human review gate.
- - AI-driven data quality, anomaly detection (data drift, schema drift, volume shifts, reconciliation breaks), automated reconciliation, and synthetic privacy-preserving test data.
- - Semantic layer and knowledge graph, plus a GPT-powered conversational interface (text-to-SQL / semantic-layer retrieval) with row- and column-level security.
- - Governance and leadership: Unity Catalog (lineage, access control, PII standards), Architecture Review Boards and AI governance forums, mentoring engineers, documentation.
MUST-HAVE SKILLS & EXPERIENCE
- - Expert-level Python, Scala and PySpark: production-ready, modular, well-tested solutions; Spark workload troubleshooting; optimizing large-scale batch and streaming pipelines using Delta Lake.
- - Strong SQL and data modelling (dimensional and normalised), schema design, data contracts.
- - Databricks expertise: Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.
- - Azure data stack: ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks), Azure Event Hubs.
- - 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, chunking and embedding strategy, vector and hybrid retrieval, prompt engineering.
- - Evaluation discipline: golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback.
- - Hands-on with LangChain, LlamaIndex or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
- - Metadata-driven frameworks: schema inference, data profiling, lineage, catalogs.
- - 12-18 years of total experience in data engineering / data platform delivery.
- - Proven enterprise-scale delivery of a medallion / lakehouse architecture.
- - Azure security and governance: Entra ID, managed identities, RBAC, POSIX ACLs, Key Vault, private endpoints, PII handling.
- - CI/CD and IaC: Azure DevOps, Terraform, Databricks Asset Bundles, automated testing of data pipelines.
- - Clear technical writing and ability to present and defend designs to engineers and non-technical stakeholders.
STRONGLY PREFERRED
- - Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modelling over a lakehouse).
- - Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale.
- - ML-based anomaly detection on time-series or transactional financial data.
- - Financial services or insurance domain (finance close, GL, subledger, reconciliation, actuarial data).
- - LLMOps / MLOps: model and prompt versioning, cost governance, observability.
- - Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.
- - dbt, Great Expectations or similar; Workday, Prism or Accounting Center exposure.
Similar jobs
- RE
Data Engineer
NewResourcesoft, Inc.
Millburn, NJ🇺🇸Hybrid18 hours agoSQLScalaETL+5Technology - AR
DataOps Engineer
NewARV-Infotech
Erie, PA🇺🇸On-site18 hours agoSQLAWSETL+5Technology - E-
Data Engineer/Senior Data Engineer/Lead/Architect
NewE-Solutions, Inc.
Irving, TX🇺🇸On-site18 hours agoSQLETLScrum+7Technology - BI
Principal Data Engineer(with Architecture experience) @w2 only
NewBURGEON IT SERVICES LLC
Houston, TX🇺🇸Hybrid18 hours agoMicroservicesSQLSQL Server+4Technology - EP
Snowflake Data Engineer With Admin Experience
NewEmpower Professionals
Fort Mill, SC🇺🇸On-site18 hours agoSQLAWSETL+10Technology - KI
Data Engineer DB2/Python
NewKanshe Infotech
Houston, TX🇺🇸On-site18 hours agoMicroservicesOracleSQL+3Technology - DI
Senior Data Engineer
NewAuto ApplyDiscord
San Francisco Bay Area🇺🇸Hybrid3 hours agoSQLLookerMachine Learning+6Technology - IC
Senior Data Engineer – Databricks & AWS- 10+ yrs- Indianapolis, United States- Onsite
NewiMedhas Consulting Services
Indianapolis, IN🇺🇸Hybrid18 hours agoSQLAWSETL+4Technology - GT
Senior Data Engineer – Cloud, Big Data & Data Platforms
NewGTSS Inc
Seattle, WA🇺🇸On-site18 hours agoDockerSQLAWS+16Technology - OP
SQL Server Data Engineer
NewOptimuss Inc.
Dallas, TX🇺🇸Hybrid18 hours agoSQLSQL ServerETL+1Technology - YS
Lead Data Engineer
NewYork Solutions, LLC
United States🇺🇸Hybrid18 hours agoAWSGitHub ActionsPython+1Technology - TA
Lead Data Engineer
NewTalentrix AI INC
Cary, NC🇺🇸Hybrid18 hours agoSQLScalaMLOps+4Technology