Haystack
← Back to Jobs
Technology

Databricks Data Engineer

Donato Technologies IncChicago, IL🇺🇸United StatesPosted 14 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Key Responsibilities
· Pipeline Development – Design, implement, and optimize scalable ETL/ELT pipelines using Databricks (PySpark, SQL, Delta Lake) to ingest structured and semi structured data from multiple sources (APIs, databases, streaming).
· Data Warehousing – Build and maintain cloud data warehouse solutions (Snowflake / Azure Synapse) – design star schemas, fact/dimension tables, and aggregate tables for high performance reporting.
· Data Modeling – Create logical and physical data models for operational and analytical use cases; implement SCD Type 2, slowly changing dimensions, and data vault methodologies where appropriate.
· Performance Tuning – Optimize Spark jobs, SQL queries, and data partitioning strategies to handle petabyte scale data with low latency.
· Governance & Quality – Implement data quality checks, monitoring, and lineage using tools like Great Expectations or custom frameworks; enforce data governance policies (GDPR/CCPA).
· Collaboration – Partner with data analysts, product managers, and engineers to translate business requirements into technical data solutions.
· CI/CD & Automation – Automate deployment of data pipelines using Azure DevOps or GitHub Actions; maintain infrastructure as code (Terraform) for data resources.

Required Skills & Experience

· Total Experience: 10+ years in data engineering or related roles.
· Cloud Data Platforms: Deep hands on experience with Databricks (notebooks, jobs, clusters, Delta Lake, Unity Catalog) – must have production level work.
· Data Warehousing: Proven experience with cloud data warehouses (Snowflake, Azure Synapse, or Redshift) – design, optimisation, and administration.
· Data Modeling: Strong knowledge of dimensional modeling (Kimball/Inmon), relational database design, and experience with tools like ER/Studio or dbt.
· Programming: Expert in Python and SQL – ability to write maintainable, production grade code.
· Big Data: Hands on with Apache Spark (PySpark), distributed computing, and performance tuning.
· Orchestration: Experience with workflow tools (Airflow, Azure Data Factory, or Prefect) for scheduling and monitoring pipelines.
· Version Control: Proficient with Git and collaborative development workflows.

Preferred Qualifications

· Experience with streaming technologies (Kafka, Event Hubs, or Kinesis).
· Knowledge of data mesh or data fabric architectures.
· Familiarity with BI tools (Power BI, Tableau, Looker).
· Databricks certification (e.g., Associate or Professional Data Engineer).
· Experience with dbt (data build tool) and transformation testing.
· Exposure to MLflow or MLOps practices.

Education & Soft Skills
· Bachelor’s or Master’s degree in Computer Science, Information Systems, or a related field (or equivalent practical experience).
· Strong communication skills
· Self starter with a problem solving mindset and ability to work independently in a hybrid environment.
Keyword: 
Skills: Digital : Snowflake~Digital : Databricks
Experience Required: 10 & Above

Skills

SQL
ETL
Looker
MLOps
MLflow
Snowflake
Tableau
Airflow
Apache
Apache Spark
Azure
Databricks
GDPR
Git
GitHub Actions
Kafka
Power BI
Python
Redshift
Terraform
Unity
Vault
dbt

Similar jobs