Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
St. Louis, MO, United States
Posted
22 hours ago
SQLAWSETLApacheApache SparkAzureDatabricksGitHub ActionsJenkinsPythonRedshiftTerraformUnityVault
Job Description
Lead Data Engineer with Databricks & AWS
Introduction:
We are seeking a Lead Data Engineer with expertise in Databricks & AWS to join our team in St. Louis, MO. The ideal candidate will play a crucial role in designing and implementing scalable, secure, and high-performance data solutions using AWS and Databricks technologies.
Responsibilities:
- Design and architect scalable, secure, and high-performance AWS and Databricks solutions
- Develop and maintain robust ETL/ELT pipelines using PySpark and Python within Databricks
- Implement medallion architecture and optimize Spark jobs for performance
- Provision and manage cloud infrastructure using Terraform for Databricks workspaces
- Write efficient SQL queries for data transformation and analytics within Databricks
- Implement and maintain CI/CD pipelines using Jenkins and GitHub Actions
- Integrate data from diverse sources into cloud storage and processing layers
- Optimize cloud costs through auto-scaling clusters and efficient resource usage
- Monitor pipeline performance, troubleshoot failures, and implement alerting and observability
Requirements:
Required Skills:
- Advanced proficiency in AWS Networking and VPC design
- Expertise in AWS Kinesis for real-time data streaming and analytics
- Deep experience with AWS Elastic Cache for distributed caching
- Extensive experience with AWS Redshift for data warehousing
- Strong knowledge of AWS PaaS services including Lambda, S3, and Glue
- Hands-on experience with AWS EKS for container orchestration
- Proficiency in AWS Cognito for identity and access management
- Expertise in Databricks, including Delta Lake, Unity Catalog, and Delta Live Tables
- Strong Python and PySpark skills for distributed data processing
- Advanced SQL skills for complex querying and optimization
- Proficiency in Terraform for infrastructure automation and management
- Experience with CI/CD tools such as Jenkins and GitHub Actions
Preferred Skills:
- Knowledge of Azure Data Lake, Data Factory, Synapse, and Key Vault
- Understanding of big data modeling and lakehouse architecture
- Experience with performance tuning in Spark and Databricks environments
- Familiarity with data governance, lineage, and compliance in cloud environments
Desired Qualifications:
- Bachelor's degree in Computer Science, Information Technology, or related field
- AWS Certified Solutions Architect – Professional
- Databricks Certified Data Engineer Professional or Apache Spark certification
Similar jobs
- ST
Data Engineer
NewStoneGate-Technologies LLC
United States🇺🇸Remote22 hours agoTechnology - TC
ETL Lead
NewTATA Consultancy Services Limited
Pittsburgh, PA🇺🇸Hybrid22 hours agoPL/SQLSQLShell+5 - TT
Lead Data Engineer
NewTDK Technologies
O'Fallon, MO🇺🇸Hybrid22 hours agoOraclePL/SQLSQL+4Technology - KI
Lead Data Engineer - Must have 15 Yrs Exp
NewKey Infotek LLC
Madison, WI🇺🇸Hybrid22 hours agoDockerSQLETL+8Technology - DT
Data Engineer-W2- Covington , KY onsite
Digipulse Technologies, Inc
Covington, KY🇺🇸On-site4 days agoSQLAWSETL+3Technology - BS
Data Engineer
NewBayOne Solutions
Oakland, CA🇺🇸On-site22 hours agoSQLAWSETL+10Technology