Haystack
← Back to Jobs
Technology

Data Engineer (W2 and Locals only)

Baanyan Software Services, Inc.Chicago, IL🇺🇸United StatesPosted 20 Jul 2026

Why This Role Stands Out

This hybrid role offers significant growth potential as you develop and optimize large-scale data processing solutions using cutting-edge technologies. You'll thrive here if you're a skilled mid-senior Data Engineer passionate about building robust data pipelines and collaborating with diverse teams to deliver impactful features. Embrace this opportunity to advance your career with a reputable company.

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

As a Senior Data Engineer  you will:

  • Develop and optimise large-scale data processing solutions using Scala, Spark, SQL, and modern data platform technologies for the components and pipelines you own.
  • Build and operate trusted data processing pipelines that handle advertiser, customer, and measurement datasets within secure cloud environments (AWS, Google Cloud Platform, Azure) and approved data-sharing ecosystems.
  • Partner with Product, Data Science, Security, Privacy, and Platform Engineering teams to deliver privacy-preserving attribution, measurement, forecasting, and analytics features.
  • Build, schedule, and maintain highly scalable batch and streaming data workflows using orchestration frameworks and cloud-native data services.
  • Implement data classification, access controls, and privacy-preserving processing techniques to ensure sensitive datasets and identifiers are handled in accordance with security and compliance requirements.
  • Contribute to the design and operation of clean-room and trusted data-sharing environments, ensuring only approved aggregated or privacy-protected outputs are made available for downstream consumption.
  • Build observability, monitoring, and operational tooling to ensure reliability, performance, and compliance of data processing platforms.
  • Troubleshoot complex data platform, performance, and pipeline issues across distributed systems.
  • Contribute to technical design, engineering best practices, and operational excellence for the components and pipelines you own.
  • Mentor junior engineers, perform thorough code reviews, and contribute to a strong engineering culture within your team.
  • Continuously improve Epsilon''s attribution, measurement, forecasting, and privacy-preserving analytics capabilities through the features and pipelines you ship.
  • Strong written and verbal English communication skills are required.
  • Good understanding of Agile/SCRUM methodologies and experience working within cross-functional product development teams.

 

Core technical skills

  • 5+ years of Data Engineering experience with strong Scala programming and hands-on Apache Spark expertise for large-scale distributed data processing on AWS and/or Google Cloud Platform.
  • Strong Python development skills for data pipelines, platform tooling, automation, and infrastructure modules .
  • Advanced SQL skills across relational databases, cloud data warehouses, and lakehouse platforms; experience handling TB-scale datasets .
  • Experience designing, building, and maintaining batch and streaming data pipelines.
  • Strong understanding of data warehousing, dimensional modelling, data quality, partitioning, and performance optimisation.
  • Experience with distributed data processing and modern lakehouse architectures (Databricks, Delta Lake, Apache Spark, or equivalent) .
  • Experience building and operating distributed data platforms at scale.
  • Experience with workflow orchestration platforms such as Airflow, Databricks Workflows, AWS Step Functions, or equivalent DAG-based systems.
  • Git or equivalent source control; unit, integration, and automated testing frameworks
  • Cloud-native development experience across AWS and/or Google Cloud Platform.
  • Strong software engineering practices including CI/CD, code reviews, observability, and production support .
  • Ability to take end-to-end ownership of features and pipelines, collaborate effectively across teams, mentor junior engineers, and deliver within tight deadlines.

Trusted environment execution skills (required)

  • Hands-on experience operating data pipelines within trusted data environments, clean rooms, secure data-sharing platforms, or equivalent privacy-preserving analytics environments
  • Experience working with sensitive datasets containing PII, customer identifiers, advertiser data, or regulated information
  • Experience applying fine-grained access controls, data governance policies, and policy-based enforcement for sensitive datasets, including PII, PCI, and other regulated-data classification tiers
  • Working knowledge of privacy-preserving data processing techniques including tokenisation, pseudonymisation, aggregation-before-export, and differential privacy concepts
  • Experience building or supporting clean-room, measurement, attribution, audience analytics, partner data-sharing, or privacy-preserving reporting solutions
  • Understanding of trust boundaries, secure data-sharing patterns, and zero-trust principles, including environments where only aggregated, anonymised, tokenised, or privacy-protected outputs may leave the trusted perimeter
  • Experience with encryption, key management, and secure handling of sensitive data using cloud-native security services
  • Experience helping build observability, monitoring, and alerting controls to detect anomalous data movement, policy violations, and potential data leakage events
  • Strong understanding of cloud-native security and governance
  •  

Skills

SQL
Scala
AWS
Encryption
Scrum
Agile
Airflow
Apache
Apache Spark
Azure
Databricks
Git
Google Cloud
Python

Similar jobs