Why This Role Stands Out
This role offers a fantastic opportunity to architect and implement cutting-edge data solutions on Google Cloud, leveraging your expertise in BigQuery and Spark to drive impactful projects. You'll thrive here if you're a seasoned data engineer passionate about building scalable pipelines and optimizing complex data workloads in a hybrid environment. Apply today to join a forward-thinking team and advance your career in data engineering.
Quick Overview
Job Description
Job Summary
We are looking for an experienced Senior Data Engineer with strong hands-on expertise in Google BigQuery, Apache Spark, PySpark, Python, and SQL. The candidate will design and develop large-scale data pipelines, optimize distributed data-processing workloads, and build scalable analytical data platforms on Google Cloud.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python, SQL, Spark, and PySpark.
- Build and optimize enterprise-scale data solutions using Google BigQuery as the analytical data warehouse.
- Develop complex BigQuery SQL involving CTEs, window functions, nested/repeated fields, ARRAY/STRUCT operations, MERGE statements, and incremental processing.
- Design efficient BigQuery tables using partitioning, clustering, materialized views, and appropriate data modeling techniques.
- Analyze BigQuery query execution plans and optimize queries to reduce slot consumption, bytes scanned, execution time, and overall processing cost.
- Implement incremental ingestion and transformation patterns rather than performing unnecessary full-table processing.
- Develop large-scale distributed processing applications using Apache Spark and PySpark.
- Work extensively with Spark DataFrames, Spark SQL, transformations, actions, joins, aggregations, and window operations.
- Troubleshoot and optimize Spark workloads using Spark UI, execution plans, DAGs, stages, tasks, and executor metrics.
- Perform advanced Spark performance tuning including partition management, repartition/coalesce strategies, predicate pushdown, partition pruning, caching/persistence, broadcast joins, and Adaptive Query Execution (AQE).
- Identify and resolve data skew, shuffle bottlenecks, executor memory issues, excessive spills, and long-running stages.
- Tune Spark configurations including executor memory, cores, shuffle partitions, serialization, and dynamic resource allocation based on workload requirements.
- Design scalable processing patterns for multi-terabyte datasets while minimizing unnecessary data movement and shuffle operations.
- Build batch and, where required, near-real-time data processing pipelines using appropriate Google Cloud Platform services.
- Integrate BigQuery and Spark with services such as Google Cloud Storage, Dataproc, Dataflow, Pub/Sub, and Cloud Composer/Airflow.
- Design dimensional and analytical data models including fact tables, dimension tables, star schemas, curated datasets, and reporting layers.
- Implement CDC, incremental loads, deduplication, late-arriving data handling, SCD Type 1/Type 2, and idempotent pipeline patterns.
- Build reusable frameworks for ingestion, transformation, validation, logging, exception handling, and pipeline monitoring.
- Implement automated data-quality checks for completeness, uniqueness, accuracy, referential integrity, schema validation, and business-rule validation.
- Troubleshoot production pipeline failures and perform root-cause analysis across Spark jobs, SQL workloads, source systems, and downstream datasets.
- Implement monitoring and alerting for pipeline failures, SLA violations, data-quality issues, and abnormal processing behavior.
- Work with structured, semi-structured, and large-volume datasets including JSON, Parquet, Avro, and CSV.
- Apply security and governance practices including IAM, service accounts, BigQuery authorized views, row-level security, column-level security, and least-privilege access.
- Participate in code reviews, technical design discussions, performance optimization, deployment, and production support.
- Collaborate with analytics, BI, data science, and application teams to deliver reliable datasets for reporting, analytics, and AI/ML use cases.
Preferred Qualifications
- 7–8+ years of overall Data Engineering experience.
- Strong production experience with BigQuery and Apache Spark/PySpark.
- Experience designing enterprise-scale cloud data platforms on Google Cloud Platform.
- Strong understanding of distributed computing, Spark internals, and query optimization.
- Experience processing datasets ranging from hundreds of gigabytes to multiple terabytes.
- Strong understanding of data warehouse architecture and dimensional modeling.
- Experience with CI/CD, Git, automated testing, and Infrastructure as Code is preferred.
- Experience supporting analytics, Looker/BI, machine learning, or AI-oriented datasets is a plus.
Core Technology Stack
BigQuery | Apache Spark | PySpark | Python | SQL | Google Cloud Platform | Dataproc | GCS | Pub/Sub | Cloud Composer | Airflow | Parquet | Avro | Git | CI/CD | Data Modeling | ETL/ELT
Similar jobs
- CG
Senior Data Engineer
NewCredit Genie
Plymouth Meeting, PA🇺🇸$1k/moRemote12 hours agoSQLSwiftAWS+6Technology - CG
Discovery Database Administrator (DBA) (Top Secret Clearance Req with Security Clearance
NewContact Government Services, LLC
Washington, DC🇺🇸$142.1k - $192.8k/yrHybridYesterdaySQLSQL ServerTechnology - CG
Senior Database Administrator II with Security Clearance
NewContact Government Services, LLC
Boston, MA🇺🇸$114.8k - $165.8k/yrHybridYesterdaySQLSQL ServerAWS+3Technology - CG
Senior Database Administrator II with Security Clearance
NewContact Government Services, LLC
Chicago, IL🇺🇸$114.8k - $165.8k/yrHybrid12 hours agoSQLSQL ServerAWS+3Technology - CG
Discovery Database Administrator (DBA) (Top Secret Clearance Req with Security Clearance
NewContact Government Services, LLC
Dallas, TX🇺🇸$142.1k - $192.8k/yrHybrid12 hours agoSQLSQL ServerTechnology - CG
Senior Database Administrator II with Security Clearance
NewContact Government Services, LLC
Chantilly, VA🇺🇸$114.8k - $165.8k/yrHybrid12 hours agoSQLSQL ServerAWS+3Technology