Haystack
← Back to Jobs
Technology

Senior Data Engineer — Databricks

MOONITSolutions Inc.United States🇺🇸United StatesPosted 7 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

We are looking for a Senior Data Engineer to join our FinTech customer team and lead the migration of custom structured finance datasets from a legacy proprietary platform onto Databricks. This is not a ticket-taking role: you will own the data models end to end deciding how complex financial instruments and their time-series behavior should be represented - and build the ingest, transform, and publish layers around them.

Much of the source system''''''''s business logic lives in a large, sparsely documented codebase rather than in any specification, so early effort goes into reading that code closely and separating deliberate financial logic from accumulated workaround. The work rewards patience over speed. You will operate independently on outcomes rather than step-by-step direction, and you will need to explain and document your decisions clearly to both engineering and business stakeholders.

Responsibilities:

  • Migrate custom structured finance datasets from the proprietary legacy system onto Databricks, recovering the business rules embedded in its codebase and validating parity through cutover.
  • Design and own the target data models, deciding how instruments and their time-series behavior are represented.
  • Build and operate the ingest, transform, and publish layers — source onboarding, quality gates, business-logic transformation, and consumer-facing publication.
  • Tune performance and cost across large-scale workloads: cluster configuration, job orchestration, query optimization, and storage layout.
  • Establish data quality, lineage, reconciliation, and observability practices appropriate to a regulated financial environment.
  • Partner with quantitative, risk, and product stakeholders to translate domain requirements into durable data structures.
  • Document models, contracts, and migration decisions to a standard that outlives the engagement.

Required Qualifications:

  • 8+ years in data engineering, with demonstrable seniority in design ownership — not just implementation.
  • Strong, proven, hands-on Databricks experience — Delta Lake, Spark (PySpark and Spark SQL), job/workflow orchestration, and performance tuning across partitioning, clustering, file sizing, and incremental processing. Production depth required; familiarity is not sufficient.
  • Demonstrated skill in data modeling and abstraction — you can explain why you structured a domain the way you did, and what you rejected.
  • Substantial experience processing and transforming financial time-series data, including the correctness concerns unique to it — point-in-time accuracy, as-of joins, bitemporal history, late-arriving and restated data, corporate actions, and calendar alignment.
  • Experience across all stages of the pipeline — ingest, transform, and publish — rather than depth in only one segment.
  • Experience with large-scale financial systems and the data volumes, auditability, and accuracy expectations that come with them.
  • Working knowledge of structured finance instruments and datasets — securitizations, ABS/MBS, loan-level and cashflow data etc.
  • Demonstrated ability to read and reason about unfamiliar production code in order to extract requirements from it. Comfort operating without documentation, and the patience to work through a large legacy system methodically rather than rewriting on assumption.
  • Strong SQL and Python; comfort with shell scripting and version-controlled, CI/CD-driven delivery.
  • Proven ability to work independently with minimal supervision and ambiguous starting requirements.

Preferred Qualifications:

  • Prior experience executing a legacy platform migration, including parity testing and cutover.
  • Familiarity with the languages and platforms common to older proprietary financial systems (e.g. legacy SQL dialects and stored procedures, C++/C#/Java monoliths, or vendor-specific scripting).
  • Databricks Unity Catalog, Delta Live Tables, and cost governance.
  • Cloud data platform experience (AWS or Azure), including object storage, IAM, and encryption/key management.
  • Orchestration tooling (Airflow, Databricks Workflows, or similar) and infrastructure-as-code.
  • Exposure to preparing datasets for downstream analytics, quantitative research, or AI/ML consumption

 

Skills

SQL
Shell
AWS
Encryption
Airflow
Azure
C#
Databricks
C++
Java
Python
Unity

Similar jobs