Haystack
← Back to Jobs
Remote
Technology
RI

Data Platform Engineer

RAVIN IT SOLUTIONS, IncUnited States🇺🇸United StatesPosted Sep 21, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
19 hours ago
SQLShellAWSETLEncryptionSnowflakeApacheApache SparkDatabricksPython

Job Description

DBA Iceberg / Lakehouse Operations Engineer (DATA PLATFORM ENGINEER)
Location: Remote
Job Summary

We are seeking a hands-on Iceberg DBA / Lakehouse Operations Engineer to manage, operate, troubleshoot, and optimize enterprise data lakehouse environments built around Apache Iceberg. The ideal candidate will have strong experience with Iceberg table administration, metadata management, performance optimization, data lifecycle operations, and production support across modern cloud data platforms.

Required Skills

  • Strong hands-on experience with Apache Iceberg and Iceberg table administration.
  • Experience supporting production Data Lakehouse environments.
  • Strong understanding of Iceberg architecture, including:
    • Catalogs
    • Snapshots
    • Manifest files
    • Metadata files
    • Partitioning
    • Schema evolution
    • Time travel
    • Table maintenance
  • Experience with Iceberg table maintenance and optimization, including:
    • Compaction / file rewrite
    • Snapshot expiration
    • Orphan file cleanup
    • Manifest optimization
    • Small-file management
  • Strong SQL and experience troubleshooting query and table performance.
  • Experience with one or more Iceberg-compatible platforms such as AWS, Snowflake, Databricks, Trino, Spark, EMR, or AWS Glue.
  • Experience with Apache Spark / PySpark for Iceberg operations and troubleshooting.
  • Strong understanding of cloud object storage such as Amazon S3, ADLS, S.
  • Experience with production monitoring, incident management, root-cause analysis, and performance troubleshooting.
  • Experience with data ingestion, ETL/ELT pipelines, and batch processing in a lakehouse environment.
  • Understanding of data security, access control, encryption, and governance in cloud data platforms.
  • Experience with automation/scripting using Python, PySpark, or Shell scripting.
  • Familiarity with CI/CD and Infrastructure-as-Code concepts is a plus.

Key Responsibilities

  • Administer and support Apache Iceberg tables and lakehouse environments in production.
  • Perform Iceberg table maintenance, optimization, compaction, and metadata cleanup.
  • Monitor table health, storage utilization, query performance, and pipeline execution.
  • Troubleshoot production issues involving Iceberg tables, Spark jobs, catalogs, metadata, partitions, and object storage.
  • Analyze and resolve small-file and data-layout issues affecting query performance.
  • Manage schema and partition evolution while maintaining data integrity.
  • Support snapshot management, time travel, rollback, and data recovery activities.
  • Optimize Iceberg tables for performance, scalability, and cost efficiency.
  • Work with data engineering teams to troubleshoot ingestion and transformation issues.
  • Develop automation for recurring operational tasks and health checks.
  • Participate in incident response, RCA, and preventive maintenance.
  • Maintain operational documentation, runbooks, and support procedures.

Similar jobs