Haystack
← Back to Jobs
Technology
AT

Senior Data Engineer (Scala, Spark & Databricks Platform)

Anagha Techno SoftNew York, NY🇺🇸United StatesPosted 25 Aug 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
New York, NY, United States
Posted
22 hours ago
SQLScalaETLSnowflakeTDDAirflowAzureBDDDatabricksGitHub ActionsPythonUnity

Job Description

Senior Data Engineer (Scala, Spark & Databricks Platform)
Role Overview
We are seeking an experienced Senior Data Engineer to design, accelerate, and scale our next-generation data platform. In this role, you will lead complex, large-scale migrations from legacy enterprise environments to advanced cloud data lakes. You will leverage spec-driven development workflows to transform high-level requirements into production-ready, idempotent data pipelines.
Operating within a robust financial services and regulatory reporting domain, you will champion modern data engineering practices, strict test-driven development (TDD), and meticulous data quality reconciliation frameworks.

Technical Requirements
Big Data & Compute (Scala / Spark)
  • Production Scala Spark : Extensive experience writing and deploying compiled Spark applications (not just notebooks).
  • Spark Optimization : Deep fluency with the DataFrame API, complex joins, window functions, custom partitioning strategy, and advanced performance tuning.
  • JVM Ecosystem : Proficiency with Gradle (including composite builds) and broader JVM tooling.
Data Ecosystem & Cloud Platform (Databricks, Snowflake & Azure)
  • Databricks Ecosystem : Hands-on experience implementing Databricks Serverless compute, governance via Unity Catalog, Databricks Asset Bundles (DABs), and automation through the Databricks CLI.
  • Snowflake Integration : Solid foundation in Snowflake schema design, performance optimization, and utilizing the Spark-Snowflake connector.
  • Microsoft Azure : Practical knowledge of Azure Data Lake Storage (ADLS), cloud networking basics, and security implementations using Entra ID, secrets, and managed identities.
Engineering Practices, Architecture & Quality
  • Modern Architecture : Practical mastery of the Medallion architecture (Bronze/Silver/Gold), building idempotent pipelines, handling schema evolution, tracing lineage, and establishing observability.
  • Testing Rigour : Committed to Test-Driven Development (TDD) using ScalaTest ( AnyFlatSpec ) and Behaviour-Driven Development (BDD) frameworks such as Concordion or equivalent.
  • Data Quality & Reconciliation : Experience creating automated parity checks against legacy systems, drift detection mechanisms, and row-level reconciliation tooling.
  • SQL Fluency : Exceptional ability to write, analyze, and reverse-engineer requirements from complex enterprise SQL scripts.
  • Secondary Tooling : Proficiency in Python for Databricks utilities and helper tooling.
Orchestration, CI/CD & Migrations
  • Workflow Automation : Advanced Airflow capabilities, including complex DAG authoring, sensors, custom retries, and strict SLA configuration.
  • CI/CD Pipelines : Experience engineering GitHub Actions pipelines featuring sharded test matrices, Artifactory integration, and secure artifact promotion across environments (Dev → QA → UAT → Prod).
  • Legacy Migrations : Proven track record of spearheading large-scale data migrations from legacy ETL infrastructure (e.g., Autosys, Informatica) to the cloud, including dependency mapping and cutover planning.

Domain & Soft Skills
  • Financial Services Domain : Strong understanding of data patterns and accuracy thresholds required for financial services and regulatory compliance reporting.
  • Spec-Driven Workflows : Proven discipline following a structured engineering lifecycle (Specifications → Technical Plans → Dissected Tasks → Implementation).

Similar jobs