Haystack
← Back to Jobs
Technology
CL

Data Engineer with Databricks-Only W2

cloudingest incUnited States🇺🇸United StatesPosted 10 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
22 hours ago
DockerSQLAWSETLMLflowSnowflakeAirflowAzureData PipelineDatabricksGitHub ActionsGoogle CloudJenkinsKafkaKubernetesLLMPythonRESTTerraformUnitydbt

Job Description

Role : Data Engineer

Location : Remote

We are seeking a highly skilled Data Engineer with handson experience in Databricks, Snowflake, and modern cloud data pipelines. The ideal candidate has deep expertise in PySpark, SQL, Delta Lake, Snowflake ELT, and distributed data processing. This role focuses on building scalable data pipelines, optimizing performance, ensuring data quality, and supporting analytics, AI/ML, and enterprise reporting workloads

Key Responsibilities
Databricks Engineering
Build and optimize ETL/ELT pipelines using PySpark, Spark SQL, Databricks Workflows, and Delta Lake.
Develop Bronze/Silver/Gold Medallion architecture pipelines.
Implement Delta Live Tables (DLT) for automated ingestion and transformation.
Manage and optimize Databricks clusters, jobs, notebooks, repos, and workflows.
Perform Spark performance tuning (partitioning, caching, AQE, broadcast joins).
Implement Unity Catalog governance (catalogs, schemas, tables, permissions).
Integrate Databricks with AWS S3 / Azure Data Lake / Kafka / APIs.
Snowflake Engineering
Design and develop Snowflake ELT pipelines using Snowflake SQL, Streams, Tasks, and Snowpipe.
Build warehouse models, fact/dimension tables, and curated datasets.
Optimize Snowflake performance (clustering, micro-partitioning, query tuning).
Implement RBAC, masking policies, row-level security, and governance.
Integrate Snowflake with Fivetran, DBT, ADF, Glue, Kafka, or custom ingestion frameworks.
Data Pipeline & Integration
Build scalable ingestion frameworks for structured, semistructured, and unstructured data.
Integrate data from databases, APIs, cloud storage, streaming sources, and enterprise systems.
Implement data quality checks, validation rules, reconciliation, and lineage.
Support ML/AI workloads, feature engineering, and model-ready datasets.
Cloud & DevOps
Work with AWS, Azure, or Google Cloud Platform cloud-native services.
Implement CI/CD using GitHub Actions, Azure DevOps, GitLab, Jenkins.
Containerize workloads using Docker and orchestrate with Kubernetes (nice to have).
Monitor pipelines using CloudWatch, Azure Monitor, Databricks metrics, PrometheGrafana.
Required Skills
Core Technical Skills
Databricks (PySpark, Spark SQL, Delta Lake, Workflows, DLT)
Snowflake (SQL, Streams, Tasks, Snowpipe, RBAC)
Python
Advanced SQL
Cloud platforms: AWS / Azure / Google Cloud Platform
Data modeling (Star schema, dimensional modeling)
ETL/ELT pipeline development
Data quality, governance, lineage
CI/CD pipelines
API integration & REST services
Nice-to-Have Skills
DBT
Kafka / Spark Structured Streaming
MLflow / Feature Store
Terraform
Airflow / ADF / Glue / Dataflow
RAG/LLM data preparation (bonus)
Healthcare, finance, or regulated industry experience

Similar jobs