Haystack
← Back to Jobs
Engineering

Java Spark Engineer

Aventine softwareBerkeley Heights, NJ🇺🇸United StatesPosted 1 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Responsibilities:

 Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)

 Lead design of batch and streaming ETL/ELT systems handling large data volumes

 Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction

 Set coding standards and lead code/design reviews across the team

 Drive technical decisions on data architecture, storage formats, and pipeline orchestration

 Mentor mid-level and junior engineers; act as a technical escalation point

 Partner with product, analytics, and platform teams to translate requirements into scalable systems

 Own production reliability — on-call ownership, incident response, root-cause analysis for pipeline failures

 Evaluate and introduce new tools/frameworks where they improve the system

 Contribute to capacity planning and cost optimization for cluster infrastructure

Required Qualifications

 Bachelor’s or Master’s degree in Computer Science, Engineering, or related field

 7+ years of professional Java development experience

 5+ years hands-on experience with Apache Spark in production environments

 Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management

 Proven track record designing systems processing terabyte+ scale data

 Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)

 Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark

 Proficiency with Kafka

 Strong grasp of CI/CD, containerization, and infrastructure-as-code practices

Similar jobs