Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Pasadena, CA, United States
Posted
4 days ago
SQLAWSETLTableauAirflowApacheApache SparkAzureData PipelineDatabricksGitGoogle CloudKafkaPower BIPythonTerraformUnity
Job Description
Role : Sr. Data Engineer
Location: Pasadena, CA
Work Arrangement: Hybrid
Job Summary
We are looking for an experienced Data Engineer with strong expertise in Databricks, PySpark, and Python to design, develop, and maintain scalable data engineering solutions. The ideal candidate will have hands-on experience building ETL/ELT pipelines, data processing frameworks, and data lake/lakehouse solutions using Databricks and cloud technologies.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Databricks, PySpark, and Python.
- Develop ETL/ELT workflows to ingest, transform, cleanse, and integrate data from multiple sources.
- Build and optimize data processing jobs using PySpark and Spark SQL.
- Work extensively with Databricks Lakehouse, Delta Lake, notebooks, workflows, and clusters.
- Develop reusable Python modules and frameworks for data processing and automation.
- Implement data quality checks, validation, error handling, and monitoring within data pipelines.
- Optimize Spark jobs, including partitioning, caching, joins, and performance tuning.
- Work with Delta Lake for data storage, transformation, versioning, and incremental processing.
- Integrate data from relational databases, APIs, files, cloud storage, and other enterprise data sources.
- Collaborate with Data Architects, Data Scientists, BI Developers, and business stakeholders to understand data requirements.
- Implement CI/CD and source-control practices for data engineering code.
- Troubleshoot production data pipeline failures and perform root cause analysis (RCA).
- Ensure data security, governance, lineage, and compliance requirements are followed.
- Participate in design discussions, code reviews, testing, deployment, and production support.
Required Skills
- Strong hands-on experience with Databricks
- Strong PySpark / Apache Spark experience
- Strong Python programming skills
- Experience developing ETL/ELT pipelines
- Strong SQL skills
- Experience with Delta Lake
- Experience with data lake/lakehouse architecture
- Experience with Spark performance tuning and optimization
- Experience working with large-volume datasets
- Strong understanding of data modeling and data engineering concepts
- Experience with Git and CI/CD
- Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform
Preferred Skills
- Databricks certification
- Experience with Azure Data Factory / AWS Glue / Airflow
- Experience with Azure Data Lake / Amazon S3
- Experience with Unity Catalog
- Experience with Kafka or other streaming technologies
- Experience with Terraform
- Experience with data governance and data quality frameworks
- Experience with Power BI, Tableau, or other BI platforms
Typical Technology Stack
Databricks | PySpark | Python | Spark SQL | Delta Lake | SQL | AWS/Azure | Data Lake | Git | CI/CD | Airflow/ADF/Glue
Similar jobs
- UL
Senior Data Engineer
NewUline, Inc.
Waukegan, IL🇺🇸$96k - $148k/yrOn-site17 hours agoSQLT-SQLETL+2Technology - DO
Director, Data Platform Engineering
Domino's
Ann Arbor, MI🇺🇸On-site3 days agoMachine LearningGenerative AIStakeholder ManagementTechnology - MI
Senior Python Data Scraping Engineer (Freelance)
NewMindrift
Austin, Texas🇺🇸$45/hrRemote2 days agoDockerAWSSelenium+4Engineering - MI
Senior Python Data Scraping Engineer (Freelance)
NewMindrift
San Antonio, Texas🇺🇸Remote2 days agoDockerAWSSelenium+4Engineering - MI
Senior Python Data Scraping Engineer (Freelance)
NewMindrift
Houston, Texas🇺🇸Remote2 days agoDockerAWSSelenium+4Engineering - MI
Senior Python Data Scraping Engineer (Freelance)
NewMindrift
Dallas, Texas🇺🇸Remote2 days agoDockerAWSSelenium+4Engineering