Haystack
← Back to Jobs
Technology
VD

Sr Google Cloud Platform Data Engineer

VDart, Inc.Atlanta, GA🇺🇸United StatesPosted Oct 10, 2026

Why This Role Stands Out

This Sr. Google Cloud Platform Data Engineer role offers a fantastic opportunity to build and operate real-time data pipelines for a massive customer data platform, leveraging cutting-edge Google Cloud technologies. If you're a seasoned data engineer with a passion for scalable solutions and impactful projects, you'll thrive in this hybrid environment. Apply today to contribute to a system used by millions and gain invaluable experience.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Atlanta, GA, United States
Posted
21 hours ago
SQLETLSnowflakeAirflowAzureBigQueryDatabricksGoogle CloudGraphQLKafkaLLMPythonRESTRedisTerraformUnityVaultgRPC

Job Description

Job Title: Sr Google Cloud Platform Data Engineer

Location: Atlanta, GA

Duration: / Term: Contract

Experience Desired: 8+ Years

Job Description:

The Data Engineer, CDP builds and operates the real-time pipelines behind Customer Data Platform — the system of record for 100M+ customer profiles, now moving to Google Cloud to deliver real-time customer context, near-real-time signals, and sub-50ms API access. Any agent, store associate, or digital channel should see a customer's prior conversation and continue from there.

You will build how clickstream, care, store, and network data flows into Cloud Spanner (low-latency store), BigQuery (analytics), and the knowledge graph behind GraphQL and MCP servers — while migrating and reusing proven ETL from the Azure estate (Databricks, Snowflake, Data Factory, Event Hubs, Cosmos DB).

Responsibilities

01. Real-Time Pipeline Engineering

  • Build and run NRT ingestion with Pub/Sub and Dataflow (Beam) into Spanner and BigQuery, with data contracts, schema evolution, exactly-once semantics, dead-letter handling, and replay
  • Implement Spanner loads and change streams for current-state customer profiles, tuning keys, batching, and hotspots to hold sub-50ms reads

02. Clickstream, CDC & Data Modeling

  • Implement reusable CDC, SCD Type 2, and late-arriving-data patterns in Beam, SQL, and PySpark; build clickstream ingestion at scale
  • Partner with NBA and T-Engage to cut signal latency to sub-five minutes; build BigQuery–Spanner sync and reverse ETL

03. Azure-to-Google Cloud Platform Migration

  • Move data from Azure to Google Cloud Platform with STS and BigQuery–Spanner sync, reusing Databricks, Data Factory, and Snowflake ETL; execute migration or retirement of Cosmos DB, Redis, Event Hubs, and Functions with dual-run, reconciliation, and rollback

04. Knowledge Graph, API & MCP Data Feeds

  • Build the pipelines that populate the customer knowledge graph on Spanner (entity and identity resolution, Customer 360) and feed the sub-50ms GraphQL API and the MCP servers on Cloud Run used by AI agents

05. Quality, Operations & Delivery

  • Implement data quality, lineage, and access controls (Dataplex, Unity Catalog, TISS-310) and IAM/VPC-SC/CMEK; ship on Cloud Run, Dataproc, Dataflow, and Composer with Terraform, CI/CD, monitoring, SLOs, and cost tuning
  • Own pipeline on-call and runbooks, join design reviews, mentor junior engineers, and partner with product, analytics, and AI/ML teams

Qualifications

  • Overall (5+ years): Bachelor's in Computer Science or related field (or equivalent); 5+ years in data engineering, 2+ building real-time pipelines in production.
  • Google Cloud data platform (3+ years): Production data pipelines on Google Cloud Platform, including IAM, VPC Service Controls, and cost awareness.
  • Streaming & NRT pipelines (3+ years): Pub/Sub, Dataflow (Beam), Kafka, Event Hubs, or Spark Structured Streaming with exactly-once semantics, watermarking, and clickstream ingestion at scale.
  • BigQuery (3+ years): Partitioning and clustering, reservations and cost control, Spanner federation, and BigQuery-to-Spanner (reverse ETL) data movement.
  • Cloud Spanner (1+ year): Schema and primary key design, indexes, hotspot avoidance, change streams, and bulk and streaming loads.
  • Data modeling & ETL (4+ years): Dimensional, Data Vault, and medallion models; idempotent ETL/ELT (CDC, SCD Type 2); expert SQL, Python, and advanced PySpark.
  • Databricks, Spark & Delta Lake (3+ years): Spark tuning, Delta Lake, Delta Live Tables, and Unity Catalog, to reuse and migrate existing ETL.
  • Azure & Snowflake (2+ years): Data Factory, Event Hubs, Cosmos DB, Redis, and Functions, plus Snowflake, to support a hybrid-cloud migration.
  • Serving & APIs (1+ year): Building data feeds for REST/gRPC or GraphQL services with p99 targets under 50ms; exposure to MCP, LLM/agent, or Vertex AI data access a plus.
  • Platform engineering (3+ years): Terraform/IaC, CI/CD, Cloud Run, Dataproc, and Composer or Airflow; observability and SLOs.
  • Testing & data quality (2+ years): Unit, integration, and data-quality testing for pipelines; reconciliation, lineage, and Dataplex or similar tools.
  • Collaboration (2+ years): Working with product, analytics, and AI/ML teams; clearly explaining design tradeoffs; mentoring junior engineers.

Key Skills:

Google Cloud Platform, Cloud Spanner, CDP, BigQuery, Pub/Sub, Snowflake, Spark

Similar jobs