Haystack
← Back to Jobs
Remote
Technology

Data Architect

Recruitment.aiUnited States🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Role: Data Architect
Location: Remote
 
Role Summary
The Data Architect owns the end-to-end architecture of the client''s Starburst/Trino data platform — from the federation and lakehouse layer through governance to the BI and semantic layer. It combines deep data-engineering  expertise (Trino/Starburst, Iceberg lakehouse, federated pipelines) with BI and semantic-layer leadership (data products, materialized views, and the reporting layer that consumes them). You set the architecture, standards, and patterns the rest of the team builds against, and translate business needs into a scalable, governed, high-performance platform.
 
Key Responsibilities
Architecture & Federation
  • Own the Starburst Enterprise / Apache Trino architecture — coordinator/worker topology, catalogs, and connectors federating diverse sources (AWS S3, Snowflake, BigQuery, Teradata, Oracle, PostgreSQL, Kafka) into a single SQL layer without heavy ETL.
  • Define the data-product and semantic-layer strategy: how curated data products and materialized views on Apache Iceberg are modeled, published, and consumed by BI tools.
Lakehouse & Modeling
  • Design and govern the lakehouse using open table formats — Apache Iceberg (Delta/Hive where relevant) — with ACID transactions, schema/partition evolution, and time travel.
  • Establish dimensional and semantic modeling standards (star/snowflake, data-mesh / data-product patterns).
Performance & Reliability
  • Set performance standards and troubleshoot at the platform level: query execution plans, distributed joins, cluster sizing/scaling, and memory allocation.
  • Define the deployment / reliability model on Kubernetes (Helm), including Dev/QA/Prod topology and release patterns.
Governance & Security
  • Architect fine-grained access control, data masking, and row-/column-level security using Starburst controls and/or Apache Ranger; define lineage, cataloging, and compliance patterns for a regulated environment.
BI & Semantic Layer Leadership
  • Own the reporting / semantic-layer architecture on Trino — how Power BI, Tableau, and Thoughtspot connect to and query the semantic layer, and how materialized views are tuned to serve them.
  • Set BI standards and quality/governance; guide the BI Developer(s) and BI Lead on modeling, performance, and self-service enablement.
Delivery & Leadership
  • Provide technical leadership and mentoring across the data and BI tracks; run architecture reviews.
  • Support UAT, releases, Dev→Prod promotion, and go-live; produce architecture documentation, run-books, and weekly status.
  • Partner with client stakeholders to translate business requirements into platform architecture and roadmap.
 
Required Skills & Qualifications
  • 10+ years in data architecture / data engineering, including senior/lead ownership of enterprise data platforms.
  • Hands-on production experience with Apache Trino / Presto or Starburst Enterprise (self-managed a strong plus).
  • Advanced SQL / ANSI SQL — window functions, CTEs, and tuning of distributed query joins.
  • Lakehouse architecture with Apache Iceberg (Delta/Hive a plus) — ACID, schema/partition evolution, materialized views.
  • Cloud data platforms — AWS and/or Google Cloud Platform (Azure a plus) and cloud object storage (S3, GCS, ADLS).
  • Data modeling — data-warehousing concepts, star/snowflake, lakehouse and data-mesh / data-product patterns.
  • Infrastructure as Code — Kubernetes (Helm), Docker, Terraform.
  • Security & governance — fine-grained access control, masking, RLS/CLS; Apache Ranger and/or Starburst controls.
  • BI & semantic layer — architecting the reporting layer on a semantic/query engine; familiarity with Power BI, Tableau, and Thoughtspot on a Trino semantic layer.
  • Programming — Python (required); Java or Scala a plus.
  • Strong communication — translating architecture into business terms.
Preferred
  • Starburst Enterprise (self-managed) and/or Trino/Presto in a regulated / financial-services environment.
  • Data Mesh implementation experience.
  • Streaming + batch integration (Kafka/MSK) alongside federated ad-hoc query.
  • Starburst certifications (Deployment / Implementation Expert).
 

Skills

Docker
Oracle
SQL
Scala
AWS
ETL
Snowflake
Tableau
Apache
Azure
BigQuery
Google Cloud
Helm
Hive
Java
Kafka
Kubernetes
PostgreSQL
Power BI
Python
REST
Terraform

Similar jobs