Quick Overview
Job Description
Data Platform Architecture & Engineering
Own the end-to-end lakehouse architecture: Bronze / Silver / Gold layer contracts, zone layout on ADLS Gen2, Delta Lake table design, partitioning, schema evolution, and retention strategy.
Design and build metadata-driven, parameterized ingestion frameworks that onboard new data sources without bespoke pipeline code for every feed.
Write canonical PySpark and Scala Spark transformation jobs that serve as the team reference; set coding and testing standards, review pull requests, and debug production incidents.
Design for scale and cost: tune Spark clusters and pools, apply partition pruning and caching strategies, and set cost guardrails as data volumes grow.
Build and automate CI/CD for Databricks and ADF pipelines in Azure DevOps using Databricks Asset Bundles and Terraform; maintain platform observability with Azure Monitor and Log Analytics.
Design repeatable patterns for batch files, database extracts, CDC feeds, and streaming ingestion using Azure Event Hubs / Kafka and Spark Structured Streaming.
AI-Augmented Ingestion & Canonical Mapping
Design and build the AI-augmented metadata ingestion framework that auto-generates bridge documents, DML statements, canonical table definitions, and control metadata from source schemas.
Build AI-assisted source-to-canonical attribute mapping: schema reasoning, data profiling, confidence scoring, and a human review / approval gate before any mapping reaches production.
Generate file-level and record-level validation rules from historical data and metadata analysis, feeding a configurable, rules-engine-backed data quality framework.
AI-Driven Data Quality, Anomaly Detection & Testing
Build AI-assisted data quality that analyses patterns across Bronze, Silver, and Gold layers to propose DQ rules beyond predefined checks.
Deliver anomaly detection covering outliers, data drift, schema drift, volume shifts, and reconciliation breaks — with actionable alerting, not noise.
Build AI-assisted automated reconciliation and test-data generation to feed the platform's automated testing framework.
Produce synthetic, privacy-preserving datasets for lower environments using differential privacy, format-preserving masking, and referential-integrity-safe generation.
Semantic Layer & Conversational Data Access
Design and build an ontology-driven semantic layer and knowledge graph modelling relationships across finance data entities, powering data discovery, semantic integration, and AI/BI tooling.
Build a GPT-powered conversational interface for natural-language querying of financial data: text-to-SQL or semantic-layer-mediated retrieval grounded in the knowledge graph, with row-level and column-level security enforced and every answer traceable to source.
Governance, Architecture Reviews & Team Leadership
Own end-to-end AI architecture decisions: model selection, RAG and retrieval design, prompt strategy, evaluation harnesses, guardrails, cost and latency budgets, and observability.
Implement Unity Catalog for cataloguing, lineage, and fine-grained access control; define PII classification, masking, tokenization, and encryption standards across every layer.
Take designs through Architecture Review Boards and AI governance forums, covering responsible AI, data residency, model approval, auditability, and human-in-the-loop controls.
Mentor data engineers, run design reviews, and set the engineering patterns the team builds on — without becoming a bottleneck.
Produce documentation and reusable components good enough for the client's team to operate the platform independently at engagement end.
MUST-HAVE SKILLS & EXPERIENCE
Programming & Data Engineering
Expert - level proficiency in Python, Scala, and PySpark, with a strong track record of designing and delivering production-ready, modular, and well-tested solutions; developing and troubleshooting Spark workloads; and optimizing large-scale batch and streaming data pipelines using Delta Lake and Spark technologies.
Strong SQL and data modelling — dimensional and normalised; schema design and data contract definition.
Databricks expertise — Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.
Azure data stack — ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks, error handling), Azure Event Hubs.
AI & Machine Learning
3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, structured output, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering.
Evaluation discipline — golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback loops; you measure AI quality, not assert it.
Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
Metadata-driven thinking — schema inference, data profiling, lineage, catalogs, and configuration-driven frameworks that onboard the next source without new code.
Architecture & Governance
12–18 years of total experience in data engineering, data platform delivery, or related disciplines.
Proven delivery of a medallion / lakehouse architecture at enterprise scale — not just familiarity with the concept.
Azure security and governance — Entra ID, managed identities, RBAC, POSIX ACLs on ADLS Gen2, Key Vault, private endpoints, and PII handling.
CI/CD and infrastructure as code — Azure DevOps, Terraform, Databricks Asset Bundles, and automated testing of data pipelines.
Clear technical writing and the ability to present and defend a design to both engineers and non-technical stakeholders.
STRONGLY PREFERRED
Knowledge graphs and ontologies: RDF/SPARQL, property graphs (Neo4j), or graph modelling over a lakehouse.
Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale, including access control and ambiguity handling.
ML-based anomaly detection on time-series or transactional financial data.
Financial services or insurance domain knowledge: finance close, general ledger, subledger, reconciliation, or actuarial data.
LLMOps and MLOps: model versioning, prompt versioning, cost governance, and observability tooling.
Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.
dbt, Great Expectations, or similar data-quality and transformation tooling.
Workday, Prism, or Accounting Center exposure.
WHAT MAKES SOMEONE SUCCESSFUL HERE
You prototype in days, not sprints — and the prototype is production-close enough to survive an architecture review.
You know where AI genuinely helps and where a deterministic rule is the better engineering answer. On finance data, that judgement matters more than enthusiasm.
You design for human review by default. Every AI-generated mapping, rule, and artefact lands in front of a reviewer with the reasoning attached.
You are comfortable working with US-based client stakeholders and can explain a technical trade-off to a finance business owner without jargon.
You leave behind documentation and patterns the client's own team can operate without you.
Similar jobs
- TE
Data Engineer - Snowflake
NewTekcel8
Atlanta, GA🇺🇸On-site18 hours agoSQLETLSnowflake+1Technology - IN
AWS Data Engineer
NewIncedo Inc
Austin, TX🇺🇸Hybrid18 hours agoDockerSQLAWS+15Technology - VB
Senior Data Engineer @ Dallas, TX
NewVbeyond Corporation
TX🇺🇸Hybrid18 hours agoSQLSnowflakeGoogle Cloud+3Technology - DW
Databricks Administrator
NewDale Workforce Solutions
New York, NY🇺🇸Hybrid18 hours agoAzureComplianceDatabricks+4 - PC
Lead Data Engineer
NewParkar Consulting Group, LLC
Dallas, TX🇺🇸On-site18 hours agoSQLAWSETL+8Technology - CO
Senior Data Operations Engineer
NewCoforge
Auburn Hills, MI🇺🇸On-site18 hours agoDockerOracleAWS+11Technology