Haystack
← Back to Jobs
Technology

Data Architect (Databricks)

Rivago infotech incNew York, NY🇺🇸United StatesPosted 12 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Role :  (Data) Databrick Architect

Location : Tarrytown New York (Onsite)

Duration: Long term Project  

Key Responsibilities & Skills:

·        Lead enterprise-scale data platform modernization initiatives, driving migration from AWS EMR, Apache NiFi, and legacy ETL frameworks to the Databricks Lakehouse Platform. 

·        Architect and implement scalable Lakehouse solutions using Databricks, Delta Lake, Unity Catalog, and Databricks Workflows. 

·        Design and govern end-to-end data pipelines for batch, streaming, CDC, and real-time data integration workloads.

·        Provide architecture leadership for large-scale AWS-based data ecosystems leveraging S3, IAM, Redshift, Glue Catalog, Airflow, and Databricks.

·        Develop enterprise data architecture standards covering data modelling, metadata management, lineage, governance, security, and compliance.

·        Drive adoption of Databricks best practices including Delta Live Tables (DLT), Auto Loader, Unity Catalog, Serverless Compute, Lakehouse Federation, and advanced optimization techniques. 

·        Lead replatforming and migration programmes involving PySpark applications, Airflow DAGs, EMR workloads, Redshift integrations, and NiFi pipelines.

·        Partner closely with business stakeholders, enterprise architects, product owners, and data science teams to define technology roadmaps and target architectures.

·        Establish governance frameworks using Unity Catalog, data quality controls, observability, monitoring, and operational excellence practices

·        Provide technical leadership, mentoring, architecture reviews, design governance, and solution sign-offs across multiple delivery teams.

·        Lead architecture workshops, executive presentations, solution assessments, and technology evaluations.

Mandatory Technical Skills:

Databricks Lakehouse Platform

Delta Lake, Delta Live Tables (DLT)

Unity Catalog

Databricks Workflows

Auto Loader

PySpark, Spark SQL, Python

AWS (S3, IAM, Glue, Redshift, EMR, Lambda)

Apache Airflow

CDC & Data Migration Frameworks

Data Governance & Security

CI/CD, GitHub, Jenkins

JD:

Lead enterprise-scale Databricks Lakehouse Architecture design and implementation on AWS.

Drive large-scale data platform modernisation and cloud transformation initiatives.

Architect scalable Medallion Architecture (Bronze, Silver, Gold) data platforms.

Lead migration of legacy EMR, NiFi, Redshift, and ETL workloads to Databricks.

Design high-performance batch and real-time data processing solutions.

Build robust ingestion frameworks using Auto Loader, Delta Lake, and Structured Streaming.

Define enterprise data governance and security standards using Unity Catalog.

Architect metadata-driven and reusable PySpark-based ETL/ELT frameworks.

Establish best practices for performance tuning, scalability, and cost optimisation.

Design and implement data quality, lineage, and observability frameworks.

Drive adoption of CI/CD, DevOps, Infrastructure as Code, and automation practices.

Collaborate with business, analytics, and engineering teams to define target-state architectures.

Conduct architecture reviews and provide technical leadership across multiple projects.

Mentor architects and senior engineers on Databricks and AWS best practices.

Design secure and scalable solutions leveraging AWS services (S3, Glue, Athena, Lambda, IAM, Redshift).

Implement and govern enterprise-wide data access, compliance, and security controls.

Evaluate and adopt latest Databricks capabilities such as DLT, Lakehouse Federation, Serverless, and MLflow.

Enable AI/ML, Generative AI, and advanced analytics use cases on the Lakehouse platform.

Create architecture roadmaps, migration strategies, standards, and governance frameworks.

Act as the primary technical advisor for customer leadership on data strategy and platform evolution.

Skills

SQL
AWS
ETL
MLflow
Airflow
Apache
Databricks
Generative AI
Jenkins
Python
Redshift
Unity

Similar jobs