Why This Role Stands Out
This hybrid role offers a fantastic opportunity to enhance your data quality and ETL/ELT testing expertise within a reputable company, working with cutting-edge technologies like Databricks and AWS. You'll thrive here if you possess strong SQL and Python skills and enjoy collaborating with diverse engineering teams to ensure data integrity. Apply now to leverage your experience and contribute to critical data pipelines.
Quick Overview
Job Description
Position: Senior Data Engineer / Databricks Engineer
Location: Dallas, TX
Duration: Long Term
Client: SWA
Experience: 7+ Years
Role Overview
We are looking for a strong Senior Data Quality Engineer / Databricks Engineer who has hands-on experience in data quality, ETL/ELT testing, Databricks, AWS, Kafka, SQL, and Python.
The consultant will be responsible for validating end-to-end data pipelines and ensuring data is accurate, reliable, complete, and production-ready. This role involves working closely with Data Engineering, Platform, Analytics, and DevOps teams.
Key Responsibilities
- Test and validate batch and streaming data pipelines for accuracy, completeness, consistency, and timeliness.
- Develop data quality checks for nulls, duplicates, schema changes, and referential integrity.
- Perform complex SQL-based data validations and verify business rules and transformations.
- Test Databricks and Apache Spark pipelines using PySpark/Scala.
- Validate ETL/ELT workflows using AWS Glue, Lambda, EMR, Step Functions, S3, Redshift, Athena, Kinesis, and DynamoDB.
- Test Kafka streaming pipelines, including data integrity, ordering, offsets, partitions, and schema changes.
- Validate Avro, JSON, and Protobuf data formats.
- Test scenarios such as duplicate events, late-arriving data, consumer failures, backfills, and reprocessing.
- Build and maintain Python-based automated data testing frameworks and reusable test utilities.
- Integrate automated data quality testing into CI/CD pipelines.
- Perform performance, regression, failover, and recovery testing for large-scale data pipelines.
- Monitor pipeline metrics, logs, and alerts using CloudWatch, Prometheus, Grafana, or similar tools.
- Troubleshoot data issues, perform root-cause analysis, and support production incidents.
- Work with teams to establish and maintain data quality SLAs/SLOs.
Required Skills
- 7+ years of experience in QA, SDET, Data Quality Engineering, or similar roles.
- Strong hands-on SQL experience for complex data validation.
- Strong Python experience for automation and data testing.
- Hands-on experience with Databricks, Apache Spark, and PySpark.
- Strong experience testing ETL/ELT and data pipelines.
- Hands-on experience with Kafka or other streaming technologies.
- Strong knowledge of AWS data services such as S3, Glue, Redshift, Lambda, and Athena.
- Experience working with large datasets and distributed systems.
- Strong debugging and analytical skill
Similar jobs
- TA
Senior Data Engineer
NewTiger Analytics Inc.
Chicago, Illinois🇺🇸RemoteYesterdaySQLAWSETL+11Technology - OT
AI Agent Data Engineer Snowflake
NewOpenmind Technologies
United States🇺🇸RemoteYesterdayETLSnowflakeAssembly+1Technology - LU
Data Engineer with Security Clearance
NewLevel Up, LLC
Chantilly, VA🇺🇸HybridYesterdaySQLSpringETL+4Technology - BI
Data Engineer Lead
NewBitwise
Richmond, VA🇺🇸HybridYesterdaySQLETLAgile+5Technology - PG
Data Engineer Matillion, BigQuery, Remote - 69795
NewPRIMUS Global Services Inc.
United States🇺🇸RemoteYesterdayMySQLSQLETL+6Technology - CR
Azure Data Engineer
NewCode Repo
United States🇺🇸RemoteYesterdaySQLETLAirflow+3Technology