Haystack
← Back to Jobs
Technology

Data Engineer (DataOps)

Mindsource IncCupertino, CA🇺🇸United StatesPosted 29 Jul 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Role: Data Engineer — DataOps

Location: Cupertino, CA (Hybrid)

Duration: Long-term

Type: Contract – W2

 

Summary:

We are seeking a capable, detail-minded Data Engineer with a DataOps focus to join our worldwide business development and strategy team. You will build and operate the data pipelines that deliver trusted sell-through, actuals, and related business data to Sales & Finance analysts — owning them from source ingestion through to the reflections and views analysts consume and keeping them reliable in production. Beyond your own pipelines, you will help harden and troubleshoot the team''s data pipeline estate as a whole.

 

This is a hands-on role for someone who can both deliver new data engineering work and independently operate, harden, and root-cause across a shared production environment. If you look forward to solving complex business problems, take pride in operational excellence, and are excited about this opportunity, please reach out to us.

 

Requirements:

  • 5+ years of data engineering (or software engineering with a strong data focus), with strong SQL.
  • Expertise in Python (Java or Scala a plus) and technologies such as Airflow, Spark, Trino/Dremio, Iceberg, Kafka, Docker.
  • Hands-on experience designing and maintaining custom ETL / data pipelines and warehouse solutions.
  • Proven ability to independently troubleshoot and root-cause production data issues — across a shared pipeline estate, not only pipelines you personally built — driving problems to their true (often upstream) cause and a durable fix, not only executing prescribed steps.
  • Strong ownership and operational discipline: rigorous separation of development and production environments, careful low-rework changes, and consistent follow-through on issues you find or create.
  • Demonstrated ownership of data quality — designing validation checks and performing root-cause analysis on data discrepancies.
  • Experience operating pipelines in production: incident response, backfills/reprocessing, deployment/release activities.
  • Ability to work beyond narrowly-scoped tasks — take an ambiguous or new problem and carry it to completion with limited oversight.
  • Familiarity with SDLC best practices, version control (Git), and CI/CD.
  • Excellent oral and written communication; able to produce clear, structured operational communication (change plans, RCAs, runbooks) and work across cross-functional teams.

 

Description:

You will build, test, and maintain the data solutions that give our Sales & Finance teams the accurate data they need to understand and adapt to changing business conditions. Most of your time is hands-on data engineering — building and supporting business data pipelines — with a meaningful, ongoing DataOps responsibility for the reliability of the shared platform.

 

Data Engineering:

  • Develop and maintain efficient, reliable methods of consuming data from a diverse set of sources with variable quality and predictability.
  • Build and enhance data products — aggregation layers, curated views, incremental-refresh logic, and reflections/VDS — using Airflow to orchestrate, schedule, and monitor workflows.
  • Own the code, business logic, transformations, and operational health (SLIs/SLOs) of your pipelines.
  • Reuse and contribute to the team''s shared libraries and utilities.
  • Understand existing solutions, fine-tune them, and support them; meet high standards on data and software quality (scope discipline, code reuse, local validation, edge-case coverage).

 

DataOps — reliability of the shared data platform, not limited to your own pipelines

  • DAG hardening & reliability across the team''s pipelines — improve resilience of the team''s data pipelines (yoursand others''): preflight cleanup, table/storage maintenance, downstream-refresh reliability, and closing gaps in failure alerting/monitoring so issues surface proactively.
  • Monitoring, data-quality checks (DQCs) & SLIs/SLOs — build monitoring pipelines, automated data-quality checks, alerting flows, and troubleshooting tooling the whole team relies on.
  • Data object & platform governance — lifetime governance (retire unused objects, eliminate references to private spaces, clean up when users leave), performance governance (identify/remediate mal-performing queries), acceleration (materialized reflections / table optimization), and routine platform/query-engine administration.
  • Impact-analysis & platform work — dataset/column-level impact analysis, platform upgrades and migrations (e.g. orchestrator and query-engine version migrations) with regression testing.
  • Production reliability & change management — own incident response and RCA for assigned areas; author clear, structured change/deployment plans and maintain runbooks and operational-readiness standards.

 

We are a rapidly growing team with plenty of interesting technical and business challenges to solve. We seek a self-starter who is willing to learn fast, adapt well to changing requirements, and work with cross-functional teams with minimal oversight.

 

Preferred Qualifications:

  • BS or MS in Engineering / Computer Science.
  • Experience with query/lakehouse engines (e.g. Dremio, Trino), Apache Iceberg maintenance (compaction, snapshot management), and Spark-based loads.
  • Experience with cloud services (AWS, Google Cloud Platform, or Azure) for data infrastructure and storage.
  • Experience with platform upgrades / migrations and pipeline reliability/hardening work in a shared codebase.
  • Familiarity with DataOps practices — automated data-quality frameworks, pipeline observability/alerting, operational-readiness gating, data lineage, and data-asset governance.
  • Experience in a Sales, Finance, or supply-chain analytics data domain (actuals, sell-through, forecasting).
  • Comfort reviewing peers'' pipeline code.

 

Skills

Docker
SQL
Scala
AWS
ETL
Airflow
Apache
Azure
Data Pipeline
Git
Google Cloud
Java
Kafka
Python

Similar jobs