Why This Role Stands Out
Leverage your expertise in Databricks and PySpark to architect and execute a cutting-edge data migration strategy, driving significant impact on a global scale. This fully remote role offers exceptional career growth and the chance to refine your skills in a highly reputable tech company, making it an ideal opportunity for experienced data engineers seeking challenging and rewarding work.
Quick Overview
Job Description
Job Title: Data Engineer with Databricks
Location: Remote
Job Type: Contract
Job Summary:
The Databricks Data Engineer owns the end-to-end migration strategy, target architecture design, and technical execution of moving legacy ETL workloads to the Databricks Lakehouse. He will establish migration standards, optimize PySpark pipelines, orchestrate complex data workflows, and deploy proprietary automation tools to ensure a seamless, high-performing transition from legacy systems.
Key Responsibilities:
Architectural Strategy & Governance
Design the target Databricks Lakehouse architecture utilizing Delta Lake, Photon, and Unity Catalog.
Establish global code refactoring standards, optimization benchmarks, and PySpark best practices.
Resolve highly complex dependency mappings and architect seamless, zero-downtime dual-run strategies.
Lead the technical deployment and integration of specialized migration accelerators.
Hands-on Engineering & Optimization
Review automated output from migration tools and manually refactor complex legacy logic into high-performing PySpark notebooks.
Eliminate legacy anti-patterns such as massive row-by-row processing and inefficient lookups.
Optimize PySpark code performance using advanced Spark features including Z-Ordering, partitioning, and caching.
Build robust Databricks Workflows and orchestrate complex DAGs based on comprehensive source lineage.
Technical Skills & Competencies:
Core Platforms: Databricks, Delta Lake, Unity Catalog, Photon, DataStage.
Languages & Frameworks: PySpark, Python, SQL, Shell Scripting.
Cloud & DevOps: AWS alongside CI/CD deployment pipelines.
Orchestration: Apache Airflow, Databricks Workflows.
Similar jobs
- BH
Snowflake Data Engineer
NewBeacon Hill
United States🇺🇸Remote23 hours agoSQLSnowflakeTechnology - CS
Sr. Data Engineer
NewCompest Solutions Inc
Detroit, MI🇺🇸On-site23 hours agoSQLAWSETL+8Technology - IR
Enterprise Data Governance Consultant
NewIndus River Technologies Inc.
Nashville, TN🇺🇸Hybrid23 hours agoComplianceHIPAAPMP+1 - BA
Lead Data Engineer
NewBooz Allen Hamilton
Herndon, VA🇺🇸$99k - $225k/yrOn-site23 hours agoDockerMongoDBShell+11Technology - SS
W2 Need - Data Engineer (Snowflake)
NewSVARA SOFTWARE SOLUTIONS INC
United States🇺🇸Hybrid23 hours agoSQLScalaSnowflake+9Technology - DS
Enterprise Data Governance Consultant
NewData Systems Integration Group
Nashville, TN🇺🇸On-site23 hours agoComplianceHIPAAPMP+1