Haystack
← Back to Jobs
Technology
IS

Data Engineer

ISOFTUnited States🇺🇸United StatesPosted Oct 5, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
21 hours ago
SQLMLOpsMachine LearningScrumAgileApacheAzureBigQueryConfluenceData PipelineDatabricksGitGitHub ActionsGoogle CloudJavaJiraKafkaPython

Job Description

Role: Data Engineer      

Location: Cincinnati, OH (Local Preferred; Open to Strong Remote Candidates)

Role: Data Engineer      

Location: Cincinnati, OH (Local Preferred; Open to Strong Remote Candidates)

Duration: 12-month contract, with a high likelihood of extension/conversion

Requirement Description: 

  • Work Authorization: Available on W2 or C2C basis

  • Interview Process:

    • Ropes Assessment 

    • Internal Screening with Account Executive 

    • Technical panel  team

  • Feedback Timeline: 24-48 hours after interview

  • Start Date: ASAP

Overview

  Isoft, a large-scale retail organization, is seeking a talented Data Engineer to join a high-performing data organization responsible for building and maintaining scalable data products that power analytics, machine learning, personalization, and real-time business operations.

  This position partners closely with Engineering, Data Science, Machine Learning, and Product teams to design, develop, and optimize modern data solutions across both batch and streaming environments. The ideal candidate combines strong data engineering expertise with solid software engineering fundamentals and a passion for building reliable, high-quality data platforms at scale.

  Success in this role requires hands-on experience with Spark-based processing, cloud-native technologies, performance optimization, automated data quality frameworks, and production-grade engineering practices.

What You'll Do:

Data Pipeline Development

  • Design, build, and maintain scalable data pipelines for ingestion, transformation, and integration across diverse data sources.

  • Develop reliable batch processing solutions using PySpark, SQL, and distributed data processing technologies.

  • Build data products that support analytics, machine learning, personalization, and operational reporting initiatives.

Data Processing & Performance Optimization

  • Optimize pipeline performance, scalability, reliability, and cost efficiency across large-scale datasets.

  • Leverage deep knowledge of Spark architecture, including partitioning, caching, shuffles, join strategies, and cluster tuning.

  • Troubleshoot and resolve production data processing bottlenecks.

Software Engineering Excellence

  • Apply strong Python and software engineering fundamentals to create maintainable, reusable, and scalable solutions.

  • Develop clean, testable code using object-oriented programming principles and engineering best practices.

  • Participate in code reviews and contribute to engineering standards across the team.

Data Quality & Reliability

  • Design and implement automated data quality frameworks and validation processes.

  • Develop testing strategies including unit, integration, regression, and end-to-end testing.

  • Ensure data accuracy, reliability, and compliance with organizational standards.

Analytics & Machine Learning Enablement

  • Partner with Data Scientists and Machine Learning Engineers to support feature engineering and machine learning workflows.

  • Help modernize and optimize core data assets supporting advanced analytics initiatives.

Collaboration & Documentation

  • Work closely with Engineering, Product, Data Science, and Machine Learning teams to deliver high-quality data solutions.

  • Create and maintain technical documentation, architectural diagrams, and operational processes.

  • Participate in Agile ceremonies, planning sessions, estimation activities, and sprint execution.

Required Qualifications :

  • 3-5+ years of experience in Data Engineering, Software Engineering, or a related field.

  • Strong hands-on experience building and supporting production data pipelines using PySpark and SQL.

  • Deep understanding of Spark architecture and performance optimization techniques, including partitions, shuffles, caching, joins, and cluster tuning.

  • Strong Python development experience.

  • Experience working with distributed data systems and large-scale datasets.

  • Experience implementing automated testing, validation, and data quality frameworks.

  • Strong understanding of data modeling concepts and distributed system fundamentals.

  • Experience with Git, GitHub, CI/CD pipelines, and modern software development practices.

  • Experience working within Agile/Scrum development environments.

  • Strong communication, collaboration, and problem-solving skills.

Preferred Qualifications:

  • Experience with Google Cloud Platform (Google Cloud Platform) and BigQuery.

  • Experience with Databricks and cloud-native data platforms.

  • Experience with Microsoft Azure or other cloud environments.

  • Experience with Kafka or other streaming technologies.

  • Familiarity with machine learning workflows and feature engineering pipelines.

  • Exposure to MLOps tools and practices.

  • Experience supporting enterprise-scale analytics and machine learning platforms.

Technical Skills:

Data Processing

  • PySpark

  • SQL

  • Spark Architecture

  • Performance Tuning

  • Distributed Data Processing

Programming

  • Python (Required)

  • Java (Preferred)

Cloud & Data Platforms

  • Google Cloud Platform (Google Cloud Platform)

  • BigQuery

  • Databricks

  • Microsoft Azure

Streaming Technologies

  • Apache Kafka or similar streaming platforms

DevOps & CI/CD

  • Git

  • GitHub

  • GitHub Actions

  • CI/CD Best Practices

Collaboration Tools

  • JIRA

  • Confluence

  • Microsoft Teams

Ideal Candidate Profile :

The ideal candidate understands not only how to use modern data processing frameworks such as Spark, but also the engineering principles that enable scalable and maintainable systems. This includes experience with:

  • Object-oriented programming

  • Data structures and algorithms

  • Memory management concepts

  • Variable scoping and application design

  • Reusable module development

  • Production-grade testing and reliability practices

You are passionate about building performant, scalable solutions while maintaining high standards for code quality, testing, and operational excellence.

Project Overview :

Initiative

Modernizing and optimizing core data assets that support advanced analytics, machine learning, and personalization initiatives.

Business Impact

You'll help build and maintain scalable data products that enable critical business capabilities, including customer analytics, machine learning, personalization, and real-time operational decision-making.

Day-to-Day Responsibilities:

  • Build and maintain scalable data pipelines using PySpark and SQL.

  • Optimize Spark workloads through partitioning, caching, shuffling, and join tuning.

  • Develop and support data products used across analytics and machine learning environments.

  • Implement automated testing and data quality frameworks.

  • Partner with Data Science, Machine Learning, Product, and Engineering teams.

  • Participate in Agile ceremonies, code reviews, and technical planning.

  • Create and maintain technical documentation.

Team & Culture

  • Collaborative, high-performing Agile environment.

  • Close partnership with Data Science, Product, Engineering, and Machine Learning teams.

  • Strong focus on engineering excellence, scalability, automation, and continuous improvement.

  • Opportunity to contribute to impactful analytics and machine learning initiatives within a major retail organization.

Why Apply?

This is an opportunity to join a modern data engineering team that is building scalable data platforms supporting advanced analytics, machine learning, and personalization at enterprise scale. You'll work with leading technologies, collaborate with cross-functional teams, and directly influence the future of data-driven decision-making.

Duration: 12-month contract, with a high likelihood of extension/conversion

Requirement Description: 

  • Work Authorization: Available on W2 or C2C basis

  • Interview Process:

    • Ropes Assessment 

    • Internal Screening with Account Executive 

    • Technical panel  team

  • Feedback Timeline: 24-48 hours after interview

  • Start Date: ASAP

Overview

  Isoft, a large-scale retail organization, is seeking a talented Data Engineer to join a high-performing data organization responsible for building and maintaining scalable data products that power analytics, machine learning, personalization, and real-time business operations.

  This position partners closely with Engineering, Data Science, Machine Learning, and Product teams to design, develop, and optimize modern data solutions across both batch and streaming environments. The ideal candidate combines strong data engineering expertise with solid software engineering fundamentals and a passion for building reliable, high-quality data platforms at scale.

  Success in this role requires hands-on experience with Spark-based processing, cloud-native technologies, performance optimization, automated data quality frameworks, and production-grade engineering practices.

What You'll Do:

Data Pipeline Development

  • Design, build, and maintain scalable data pipelines for ingestion, transformation, and integration across diverse data sources.

  • Develop reliable batch processing solutions using PySpark, SQL, and distributed data processing technologies.

  • Build data products that support analytics, machine learning, personalization, and operational reporting initiatives.

Data Processing & Performance Optimization

  • Optimize pipeline performance, scalability, reliability, and cost efficiency across large-scale datasets.

  • Leverage deep knowledge of Spark architecture, including partitioning, caching, shuffles, join strategies, and cluster tuning.

  • Troubleshoot and resolve production data processing bottlenecks.

Software Engineering Excellence

  • Apply strong Python and software engineering fundamentals to create maintainable, reusable, and scalable solutions.

  • Develop clean, testable code using object-oriented programming principles and engineering best practices.

  • Participate in code reviews and contribute to engineering standards across the team.

Data Quality & Reliability

  • Design and implement automated data quality frameworks and validation processes.

  • Develop testing strategies including unit, integration, regression, and end-to-end testing.

  • Ensure data accuracy, reliability, and compliance with organizational standards.

Analytics & Machine Learning Enablement

  • Partner with Data Scientists and Machine Learning Engineers to support feature engineering and machine learning workflows.

  • Help modernize and optimize core data assets supporting advanced analytics initiatives.

Collaboration & Documentation

  • Work closely with Engineering, Product, Data Science, and Machine Learning teams to deliver high-quality data solutions.

  • Create and maintain technical documentation, architectural diagrams, and operational processes.

  • Participate in Agile ceremonies, planning sessions, estimation activities, and sprint execution.

Required Qualifications :

  • 3-5+ years of experience in Data Engineering, Software Engineering, or a related field.

  • Strong hands-on experience building and supporting production data pipelines using PySpark and SQL.

  • Deep understanding of Spark architecture and performance optimization techniques, including partitions, shuffles, caching, joins, and cluster tuning.

  • Strong Python development experience.

  • Experience working with distributed data systems and large-scale datasets.

  • Experience implementing automated testing, validation, and data quality frameworks.

  • Strong understanding of data modeling concepts and distributed system fundamentals.

  • Experience with Git, GitHub, CI/CD pipelines, and modern software development practices.

  • Experience working within Agile/Scrum development environments.

  • Strong communication, collaboration, and problem-solving skills.

Preferred Qualifications:

  • Experience with Google Cloud Platform (Google Cloud Platform) and BigQuery.

  • Experience with Databricks and cloud-native data platforms.

  • Experience with Microsoft Azure or other cloud environments.

  • Experience with Kafka or other streaming technologies.

  • Familiarity with machine learning workflows and feature engineering pipelines.

  • Exposure to MLOps tools and practices.

  • Experience supporting enterprise-scale analytics and machine learning platforms.

Technical Skills:

Data Processing

  • PySpark

  • SQL

  • Spark Architecture

  • Performance Tuning

  • Distributed Data Processing

Programming

  • Python (Required)

  • Java (Preferred)

Cloud & Data Platforms

  • Google Cloud Platform (Google Cloud Platform)

  • BigQuery

  • Databricks

  • Microsoft Azure

Streaming Technologies

  • Apache Kafka or similar streaming platforms

DevOps & CI/CD

  • Git

  • GitHub

  • GitHub Actions

  • CI/CD Best Practices

Collaboration Tools

  • JIRA

  • Confluence

  • Microsoft Teams

Ideal Candidate Profile :

The ideal candidate understands not only how to use modern data processing frameworks such as Spark, but also the engineering principles that enable scalable and maintainable systems. This includes experience with:

  • Object-oriented programming

  • Data structures and algorithms

  • Memory management concepts

  • Variable scoping and application design

  • Reusable module development

  • Production-grade testing and reliability practices

You are passionate about building performant, scalable solutions while maintaining high standards for code quality, testing, and operational excellence.

Project Overview :

Initiative

Modernizing and optimizing core data assets that support advanced analytics, machine learning, and personalization initiatives.

Business Impact

You'll help build and maintain scalable data products that enable critical business capabilities, including customer analytics, machine learning, personalization, and real-time operational decision-making.

Day-to-Day Responsibilities:

  • Build and maintain scalable data pipelines using PySpark and SQL.

  • Optimize Spark workloads through partitioning, caching, shuffling, and join tuning.

  • Develop and support data products used across analytics and machine learning environments.

  • Implement automated testing and data quality frameworks.

  • Partner with Data Science, Machine Learning, Product, and Engineering teams.

  • Participate in Agile ceremonies, code reviews, and technical planning.

  • Create and maintain technical documentation.

Team & Culture

  • Collaborative, high-performing Agile environment.

  • Close partnership with Data Science, Product, Engineering, and Machine Learning teams.

  • Strong focus on engineering excellence, scalability, automation, and continuous improvement.

  • Opportunity to contribute to impactful analytics and machine learning initiatives within a major retail organization.

Why Apply?

This is an opportunity to join a modern data engineering team that is building scalable data platforms supporting advanced analytics, machine learning, and personalization at enterprise scale. You'll work with leading technologies, collaborate with cross-functional teams, and directly influence the future of data-driven decision-making.

Similar jobs