Haystack
← Back to Jobs
Remote
Technology

Data Engineer

Compunnel Inc.Aliso Viejo, CA🇺🇸United StatesPosted 18 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Job Summary: We are seeking an experienced Data Engineer with strong expertise in cloud-based data engineering, ETL/ELT development, distributed data processing, and data quality. The ideal candidate will have hands-on experience with major cloud platforms, strong Python and SQL skills, and experience building and optimizing scalable data pipelines. Exposure to multiple cloud environments, DevOps/DataOps practices, metadata-driven frameworks, and cloud cost optimization is preferred.

Key Responsibilities: Design, develop, maintain, and optimize scalable ETL/ELT data pipelines for large-scale data processing. Develop data engineering solutions using Python and SQL across cloud-based data platforms. Build and optimize data processing workflows using distributed processing frameworks such as Apache Spark and Databricks. Develop and manage workflow orchestration and scheduling using tools such as Apache Airflow, Cloud Composer, and Azure Data Factory. Work with cloud-native data services such as BigQuery, Snowflake, PostgreSQL, and Cloud SQL.

Implement data quality, governance, security, and validation processes across data pipelines and platforms. Troubleshoot data pipelines, investigate defects, and implement effective resolutions using cloud data platforms. Develop and maintain metadata-driven data frameworks and solutions. Implement CI/CD pipelines and support DevOps/DataOps practices for data engineering workflows. Monitor and optimize data pipelines for performance, reliability, scalability, and cost efficiency.

Collaborate with data engineers, software engineers, business stakeholders, and other technical teams to deliver data solutions. Identify opportunities for automation and optimization using AI platforms and modern data engineering technologies. Support cloud cost optimization initiatives across data platforms and workloads. Maintain technical documentation and communicate technical findings and recommendations to stakeholders.

Required Qualifications: 58+ years of overall experience in data engineering, software engineering, or related roles. 35+ years of hands-on experience with at least one major cloud platform such as Google Cloud Platform, AWS, or Azure. 12+ years of exposure to a second cloud environment. 46+ years of experience developing and optimizing ETL/ELT pipelines for large-scale data processing. 35+ years of strong programming experience with Python and SQL. 24+ years of experience with distributed data processing frameworks such as Apache Spark or Databricks. 24+ years of experience with workflow orchestration and scheduling tools such as Apache Airflow, Cloud Composer, or Azure Data Factory. 23+ years of experience working with cloud-native data services such as BigQuery, Snowflake, PostgreSQL, or Cloud SQL. 23+ years of experience implementing data quality, governance, and security best practices. 12+ years of experience with CI/CD pipelines and DevOps/DataOps practices. 35+ years of experience troubleshooting and resolving defects using cloud data platforms. 23+ years of experience working with metadata-driven frameworks.

Strong verbal, written, presentation, and communication skills. Proficiency with Microsoft Office Suite.

Preferred Qualifications: Cloud certification, preferably a Google Cloud Platform certification. Hands-on experience with PySpark, Dataproc, Cloud Composer, and Cloud Functions. Experience with Scala or Java. Experience with AI platforms for data engineering optimization and automation. Experience with cloud cost optimization. Experience working across multiple cloud environments.

Notes: Mandatory Areas Must Have Skills Data Engineer Primary Skillset(s) / Experience Needed 58+ years of overall experience in data engineering, software engineering, or related roles 35+ years of hands-on experience with at least one major cloud platform (Google Cloud Platform, AWS, or Azure), with 12+ years exposure to a second cloud environment 46+ years of experience developing and optimizing ETL/ELT pipelines for large-scale data processing 35+ years of strong programming experience in Python and SQL (Scala or Java is a plus) 24+ years of experience with distributed data processing frameworks such as Apache Spark Domain Experience (If any ) Healthcare experience Must have Certifications None Location - Remote Onsite Requirement - Remote Number of days onsite - Remote Education: Bachelors Degree

Skills

SQL
Scala
AWS
ETL
Snowflake
Airflow
Apache
Apache Spark
Azure
BigQuery
Databricks
Google Cloud
Java
Microsoft Office
PostgreSQL
Python

Similar jobs