Haystack
← Back to Jobs
Technology

W2 Role : Data Platform Engineer

Wise Skulls Corp.Manor, TX🇺🇸United StatesPosted 12 Aug 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to build and operate a cutting-edge cloud-based data lakehouse, empowering data professionals with reliable, high-performance data. You'll thrive here if you possess deep distributed systems and data engineering expertise coupled with strong software engineering practices, and you'll enjoy the flexibility of a hybrid work environment in either Austin or Sunnyvale. Apply now to contribute to a reputable company and grow your skills in a dynamic tech landscape.

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

There are 2 locations: Austin, TX or Sunnyvale, CA (Hybrid)

Job Summary
  We are seeking an experienced Data Platform Engineer to architect, build, and operate our cloud-based data lakehouse platform. The role is responsible for developing scalable
  data infrastructure and processing pipelines built on open table formats, distributed query and compute engines, and containerized cloud infrastructure — enabling analysts, data
  scientists, and downstream applications to access reliable, high-performance data at scale. The ideal candidate combines deep distributed-systems and data engineering expertise with
  strong software engineering practices and a passion for building robust, efficient platforms.
 
  Key Responsibilities
  Lakehouse Architecture (Apache Iceberg)
  - Architect, implement, and maintain the data lakehouse using Apache Iceberg table formats.
  - Manage table lifecycle including partitioning, schema evolution, snapshot and metadata management, and compaction.
  - Optimize table layout and file sizing for query performance, storage efficiency, and cost.
  - Implement data governance, time-travel, and ACID transaction patterns across the lakehouse.
  
  Distributed Query & Compute (Trino & Spark)
  - Deploy, tune, and operate Trino clusters for interactive and federated analytics; manage connectors, catalogs, and resource groups.
  - Build and optimize large-scale batch and streaming data processing with Apache Spark.
  - Troubleshoot query plans, job performance, spills, and shuffle bottlenecks across distributed workloads.
  - Balance capacity, concurrency, and cost across query and compute engines.
  
  Data Engineering (Scala & Python)
  - Develop scalable data pipelines, frameworks, and services in Scala and Python.
  - Build reusable ingestion, transformation, and data-quality libraries for batch and streaming.
  - Write clean, well-tested, performant code following software engineering best practices.
  - Integrate pipelines with the Iceberg lakehouse and Trino/Spark processing layers.
  
  Cloud & Kubernetes Infrastructure
  - Design and operate data platform components on Kubernetes (EKS, AKS, GKE, or self-managed).
  - Build and manage containerized workloads, Helm charts, and autoscaling for data services.
  - Provision and manage cloud infrastructure (AWS, Azure, or Google Cloud Platform), including object storage and compute.
  - Implement infrastructure as code, CI/CD, monitoring, and observability for the platform. 
 
  Platform Reliability & Collaboration
  - Ensure the reliability, scalability, and security of the data platform across environments.
  - Automate deployment, testing, and operational workflows to eliminate toil.
  - Partner with data consumers to understand requirements and improve data accessibility.
  - Mentor engineers and promote best practices in data and platform engineering.

Skills

Scala
AWS
Apache
Apache Spark
Azure
Google Cloud
Helm
Kubernetes
Python

Similar jobs