Haystack
← Back to Jobs
Technology

Data Platform Engineer

Wise Skulls Corp.Manor, TX🇺🇸United StatesPosted 10 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Data Platform Engineer

Location: Austin, TX or Sunnyvale, CA — Hybrid
Duration: 6+ Months Contract, with possibility of extension

Job Summary

We are seeking an experienced Data Platform Engineer to architect, build, and operate a scalable cloud-based data lakehouse platform. The ideal candidate will have strong hands-on experience with Apache Iceberg, Trino, Apache Spark, Scala, Python, Kubernetes, and cloud infrastructure.

This role will focus on building reliable and high-performance data platforms that support analytics, data science, and downstream applications at scale.

Key Responsibilities

  • Design, implement, and maintain a cloud-based data lakehouse using Apache Iceberg.
  • Manage Iceberg table lifecycle, including partitioning, schema evolution, snapshots, metadata management, compaction, time travel, and ACID transactions.
  • Optimize table layouts, file sizes, query performance, storage efficiency, and cloud costs.
  • Deploy, configure, tune, and operate Trino clusters for interactive and federated analytics.
  • Manage Trino connectors, catalogs, resource groups, concurrency, and workload performance.
  • Develop and optimize large-scale batch and streaming workloads using Apache Spark.
  • Troubleshoot distributed query and processing issues, including query plans, spills, shuffle bottlenecks, and job performance.
  • Develop scalable data pipelines and reusable frameworks using Scala and Python.
  • Build ingestion, transformation, and data-quality frameworks for batch and streaming environments.
  • Integrate data pipelines with Apache Iceberg, Trino, and Spark.
  • Design and operate data platform components on Kubernetes, including EKS, AKS, GKE, or self-managed Kubernetes.
  • Build and manage Docker/containerized workloads, Helm charts, and autoscaling.
  • Provision and manage cloud infrastructure across AWS, Azure, or Google Cloud Platform, including object storage and compute resources.
  • Implement Infrastructure as Code (IaC), CI/CD, monitoring, logging, and observability.
  • Ensure platform reliability, scalability, security, and operational efficiency.
  • Automate deployment, testing, and operational processes to reduce manual effort.
  • Collaborate with data engineers, data scientists, analysts, and application teams to improve data accessibility and platform capabilities.
  • Mentor engineers and promote data platform and software engineering best practices.

Required Skills

  • Strong hands-on experience with Apache Iceberg and lakehouse architecture.
  • Strong experience with Trino and distributed query processing.
  • Strong experience with Apache Spark.
  • Strong programming experience in Scala and Python.
  • Hands-on experience with Kubernetes and containerized workloads.
  • Experience with at least one major cloud platform: AWS, Azure, or Google Cloud Platform.
  • Experience with cloud object storage and distributed data platforms.
  • Experience with Helm, CI/CD, Infrastructure as Code, monitoring, and observability.
  • Strong understanding of distributed systems, data engineering, and platform engineering.
  • Experience building scalable, reliable, and high-performance data pipelines.

Preferred Qualifications

  • Experience operating production-scale lakehouse platforms.
  • Experience with streaming data architectures.
  • Experience with data governance, security, and access-control frameworks.
  • Experience optimizing distributed workloads for performance and cloud cost.
  • Strong software engineering fundamentals, including testing, code quality, and automation.

Work Arrangement: Hybrid — Austin, TX or Sunnyvale, CA

Note: Candidates must be able to work in the required hybrid location.

Skills

Docker
Scala
AWS
Apache
Apache Spark
Azure
Google Cloud
Helm
Kubernetes
Python

Similar jobs