Haystack
← Back to Jobs
Technology
CS

Sr Data Engineer

Celer Soft LLCMinneapolis, MN🇺🇸United StatesPosted 19 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Senior Data Engineer

Locations: Minneapolis, MN / Plano, TX Onsite

Experience: 10+ years



Role: Senior Data Engineer - Databricks, Python, PySpark

Job Summary

We are seeking a highly experienced Senior Data Engineer with 10+ years of expertise in designing, developing, and optimizing enterprise-scale data platforms and pipelines. The ideal candidate will have strong hands-on experience with Databricks, Python, PySpark, Spark SQL, Delta Lake, and cloud-based data engineering solutions.

The candidate will work closely with data architects, business stakeholders, analytics teams, and application teams to build scalable, reliable, and high-performance batch and streaming data solutions.

Responsibilities

  • Design, develop, test, deploy, and maintain scalable ETL/ELT data pipelines using Databricks, Python, and PySpark.

  • Build data ingestion and transformation workflows for structured, semi-structured, and unstructured data.

  • Develop and maintain Databricks notebooks, workflows, jobs, and reusable data engineering frameworks.

  • Implement Delta Lake solutions, including ACID transactions, schema enforcement, schema evolution, and version management.

  • Work with Spark SQL and advanced SQL to perform complex data transformations and aggregations.

  • Optimize Spark applications by addressing data skew, partitioning, caching, joins, shuffle operations, and cluster configuration.

  • Design and implement data pipelines following Medallion Architecture principles, including Bronze, Silver, and Gold layers.

  • Develop batch and near-real-time data processing solutions.

  • Implement data quality checks, validation rules, error handling, logging, monitoring, and alerting.

  • Support data migration and modernization initiatives from legacy platforms to cloud-based Databricks environments.

  • Collaborate with architects and stakeholders to define data models, schemas, interfaces, and data contracts.

  • Troubleshoot production issues, perform root-cause analysis, and implement permanent corrective actions.

  • Develop unit tests, integration tests, and automated validation processes for data pipelines.

  • Follow best practices for source control, CI/CD, code reviews, documentation, and deployment management.

  • Support cloud security, governance, access control, data lineage, and compliance requirements.

  • Participate in Agile ceremonies, sprint planning, technical discussions, and release activities.

  • Mentor junior and mid-level engineers and provide technical guidance to the team.


Required Skills

  • 10+ years of experience in data engineering, software engineering, or a related field.

  • Strong hands-on experience with Databricks.

  • Advanced programming experience with Python.

  • Strong expertise in PySpark and Apache Spark.

  • Advanced SQL skills, including query optimization and complex data transformations.

  • Experience developing Databricks notebooks, jobs, workflows, and cluster configurations.

  • Strong experience with Delta Lake and lakehouse architecture.

  • Experience designing scalable batch and streaming data pipelines.

  • Knowledge of Medallion Architecture and enterprise data modeling.

  • Experience with data quality, validation, monitoring, and production support.

  • Strong understanding of performance tuning and optimization for large-volume data processing.

  • Experience with Git-based version control and CI/CD practices.

  • Strong communication, troubleshooting, analytical, and collaboration skills.

  • Ability to work onsite in Minneapolis or Plano.


Preferred Skills

  • Experience with Azure Databricks and Microsoft Azure services.

  • Experience with Azure Data Factory, ADLS Gen2, Azure Event Hubs, or Azure Functions.

  • Experience with AWS or Google Cloud Platform data services.

  • Knowledge of Unity Catalog, data governance, RBAC, and data lineage.

  • Experience with Databricks Auto Loader, Delta Live Tables, and Databricks Asset Bundles.

  • Experience with Apache Kafka or other messaging technologies.

  • Experience with Airflow or other workflow orchestration tools.

  • Knowledge of Terraform and infrastructure-as-code practices.

  • Experience with Snowflake, Synapse, or other cloud data warehouses.

  • Familiarity with Docker, Kubernetes, Jenkins, GitHub Actions, or Azure DevOps.

  • Experience supporting BI, analytics, machine learning, or AI data platforms.


Education

  • Bachelor s degree in Computer Science, Information Technology, Engineering, or a related field preferred.

  • Equivalent professional experience may be considered.


Skills

Docker
SQL
AWS
ETL
Machine Learning
Snowflake
Agile
Airflow
Apache
Apache Spark
Azure
Databricks
Git
GitHub Actions
Google Cloud
Jenkins
Kafka
Kubernetes
Python
Terraform
Unity

Similar jobs