Haystack
← Back to Jobs
Technology
PE

Senior Big Query & Spark Data Engineer

Peritus Inc.United States🇺🇸United StatesPosted 10 Sept 2026

Why This Role Stands Out

This role offers a fantastic opportunity to architect and implement cutting-edge data solutions on Google Cloud, leveraging your expertise in BigQuery and Spark to drive impactful projects. You'll thrive here if you're a seasoned data engineer passionate about building scalable pipelines and optimizing complex data workloads in a hybrid environment. Apply today to join a forward-thinking team and advance your career in data engineering.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
6 days ago
SQLETLLookerMachine LearningAirflowApacheApache SparkBigQueryGitGoogle CloudPython

Job Description

Job Summary

We are looking for an experienced Senior Data Engineer with strong hands-on expertise in Google BigQuery, Apache Spark, PySpark, Python, and SQL. The candidate will design and develop large-scale data pipelines, optimize distributed data-processing workloads, and build scalable analytical data platforms on Google Cloud.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT data pipelines using Python, SQL, Spark, and PySpark.
  • Build and optimize enterprise-scale data solutions using Google BigQuery as the analytical data warehouse.
  • Develop complex BigQuery SQL involving CTEs, window functions, nested/repeated fields, ARRAY/STRUCT operations, MERGE statements, and incremental processing.
  • Design efficient BigQuery tables using partitioning, clustering, materialized views, and appropriate data modeling techniques.
  • Analyze BigQuery query execution plans and optimize queries to reduce slot consumption, bytes scanned, execution time, and overall processing cost.
  • Implement incremental ingestion and transformation patterns rather than performing unnecessary full-table processing.
  • Develop large-scale distributed processing applications using Apache Spark and PySpark.
  • Work extensively with Spark DataFrames, Spark SQL, transformations, actions, joins, aggregations, and window operations.
  • Troubleshoot and optimize Spark workloads using Spark UI, execution plans, DAGs, stages, tasks, and executor metrics.
  • Perform advanced Spark performance tuning including partition management, repartition/coalesce strategies, predicate pushdown, partition pruning, caching/persistence, broadcast joins, and Adaptive Query Execution (AQE).
  • Identify and resolve data skew, shuffle bottlenecks, executor memory issues, excessive spills, and long-running stages.
  • Tune Spark configurations including executor memory, cores, shuffle partitions, serialization, and dynamic resource allocation based on workload requirements.
  • Design scalable processing patterns for multi-terabyte datasets while minimizing unnecessary data movement and shuffle operations.
  • Build batch and, where required, near-real-time data processing pipelines using appropriate Google Cloud Platform services.
  • Integrate BigQuery and Spark with services such as Google Cloud Storage, Dataproc, Dataflow, Pub/Sub, and Cloud Composer/Airflow.
  • Design dimensional and analytical data models including fact tables, dimension tables, star schemas, curated datasets, and reporting layers.
  • Implement CDC, incremental loads, deduplication, late-arriving data handling, SCD Type 1/Type 2, and idempotent pipeline patterns.
  • Build reusable frameworks for ingestion, transformation, validation, logging, exception handling, and pipeline monitoring.
  • Implement automated data-quality checks for completeness, uniqueness, accuracy, referential integrity, schema validation, and business-rule validation.
  • Troubleshoot production pipeline failures and perform root-cause analysis across Spark jobs, SQL workloads, source systems, and downstream datasets.
  • Implement monitoring and alerting for pipeline failures, SLA violations, data-quality issues, and abnormal processing behavior.
  • Work with structured, semi-structured, and large-volume datasets including JSON, Parquet, Avro, and CSV.
  • Apply security and governance practices including IAM, service accounts, BigQuery authorized views, row-level security, column-level security, and least-privilege access.
  • Participate in code reviews, technical design discussions, performance optimization, deployment, and production support.
  • Collaborate with analytics, BI, data science, and application teams to deliver reliable datasets for reporting, analytics, and AI/ML use cases.

Preferred Qualifications

  • 7–8+ years of overall Data Engineering experience.
  • Strong production experience with BigQuery and Apache Spark/PySpark.
  • Experience designing enterprise-scale cloud data platforms on Google Cloud Platform.
  • Strong understanding of distributed computing, Spark internals, and query optimization.
  • Experience processing datasets ranging from hundreds of gigabytes to multiple terabytes.
  • Strong understanding of data warehouse architecture and dimensional modeling.
  • Experience with CI/CD, Git, automated testing, and Infrastructure as Code is preferred.
  • Experience supporting analytics, Looker/BI, machine learning, or AI-oriented datasets is a plus.

Core Technology Stack

BigQuery | Apache Spark | PySpark | Python | SQL | Google Cloud Platform | Dataproc | GCS | Pub/Sub | Cloud Composer | Airflow | Parquet | Avro | Git | CI/CD | Data Modeling | ETL/ELT

Similar jobs