Haystack
← Back to Jobs
Engineering

Senior Databricks Engineer

Purple Drive Technologies LLCMalvern, PA🇺🇸United StatesPosted 21 Jul 2026

Why This Role Stands Out

This hybrid role offers an exciting opportunity to leverage your deep expertise in Databricks and Apache Spark to build cutting-edge data solutions and explore Generative AI. You'll thrive here if you're a seasoned data engineer passionate about scalable pipelines, robust governance, and innovative MLOps practices. Apply now to join a forward-thinking team and make a significant impact.

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Job Title: Senior Databricks Engineer

Location: Malvern, PA

Job Type: Full-Time

Experience: 10-15 Years



Job Summary

We are seeking a highly experienced Senior Databricks Engineer with deep expertise in Apache Spark, Databricks Lakehouse Platform, Delta Lake, AWS, and modern data engineering. The ideal candidate will design and build scalable, secure, and high-performance data pipelines while implementing enterprise-grade data governance, CI/CD, and MLOps practices. Experience with Generative AI, LLMs, and Vector Search is highly desirable.



Required Skills

Databricks & Apache Spark

  • 10-15 years of experience in Data Engineering
  • Expert in:
    • Databricks Lakehouse Platform
    • Apache Spark
    • Spark SQL
    • PySpark
    • DataFrame API
    • Distributed Data Processing
    • Spark Performance Tuning
    • Query Optimization

Delta Lake & Lakehouse Architecture

  • Delta Lake
  • Medallion Architecture (Bronze, Silver, Gold)
  • Delta Live Tables (DLT)
  • ACID Transactions
  • Z-Ordering
  • Data Optimization
  • Time Travel
  • Schema Evolution

Data Pipelines & Orchestration

  • ETL / ELT Development
  • Databricks Workflows
  • Databricks Jobs
  • Delta Live Tables (DLT)
  • Workflow Automation
  • Batch Processing
  • Streaming Data Pipelines
  • Error Handling
  • Retry Mechanisms

Programming Languages

  • Python
  • PySpark
  • SQL
  • Scala (Preferred)

AWS Cloud & Storage

  • Amazon S3
  • Amazon Redshift
  • AWS Glue
  • Amazon Kinesis
  • AWS Step Functions
  • AWS IAM
  • AWS KMS
  • Cross-Account IAM Roles
  • S3 Bucket Policies

Databricks Governance & Security

  • Unity Catalog
  • Data Lineage
  • Row-Level Security
  • Column-Level Security
  • Data Governance
  • Secure Data Sharing
  • Compliance

Compute & Cost Optimization

  • Cluster Policies
  • Instance Profiles
  • Spot Instances
  • Compute Optimization
  • Cost Management

CI/CD & DevOps

  • Git
  • Databricks Git Folders
  • CI/CD Pipelines
  • Version Control
  • Deployment Automation

AI & MLOps

  • MLflow
  • Model Registry
  • Experiment Tracking
  • Feature Engineering
  • Large Language Models (LLMs)
  • Vector Search
  • Generative AI Integrations



Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines using Databricks, Apache Spark, and Delta Lake
  • Build and optimize batch and streaming data pipelines processing data from multiple enterprise sources
  • Develop high-performance PySpark and Spark SQL solutions, optimizing distributed processing and execution plans
  • Implement Medallion Architecture (Bronze, Silver, Gold) using Delta Lake best practices
  • Design and automate resilient workflows using Databricks Workflows, Jobs, and Delta Live Tables (DLT)
  • Configure Unity Catalog for enterprise data governance, lineage, and fine-grained access controls
  • Integrate Databricks with AWS services including Amazon S3, Redshift, Glue, Kinesis, and Step Functions
  • Implement secure cloud architectures using IAM roles, S3 bucket policies, and AWS KMS customer-managed encryption keys
  • Optimize Databricks compute resources through cluster policies, instance profiles, and Spot Instances
  • Collaborate with Data Scientists, ML Engineers, Analysts, and BI teams to support analytics, machine learning, and Generative AI initiatives
  • Implement CI/CD pipelines, Git-based development workflows, and deployment automation
  • Utilize MLflow for experiment tracking, model management, and MLOps lifecycle support



Preferred Qualifications

  • Experience with real-time streaming architectures and event-driven data platforms
  • Hands-on experience with Generative AI, LLMs, Vector Databases, and Retrieval-Augmented Generation (RAG)
  • Databricks Certified Data Engineer Professional or Associate
  • AWS Certified Data Analytics or Solutions Architect certification
  • Experience with Terraform or Infrastructure-as-Code (IaC)
  • Experience supporting enterprise-scale cloud data platforms

Skills

Unity

Similar jobs