Haystack
← Back to Jobs
Technology
SI

Data Engineer

Sabio infotechTorrance, CA🇺🇸United StatesPosted Sep 21, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Torrance, CA, United States
Posted
Yesterday
ScalaAWSETLAirflowApacheApache SparkData PipelineHadoopJavaPythonRedshift

Job Description

Position: Data Engineer

Location: Torrance, California (3 Days onsite)

Employment Type: Full-Time

Responsibilities

Develop and Maintain Data Integration Solutions:

· Design and implement data integration workflows using AWS Glue/EMR,AWS MWAA(Airflow), Lambda, Redshift

· Demonstrate proficiency in Pyspark, Apache Spark and Python for data processing large datasets

· Ensure data is accurately and efficiently extracted, transformed, and loaded into target systems.

Ensure Data Quality and Integrity:

· Validate and cleanse data to maintain high data quality.

· Ensure data quality and integrity by implementing monitoring, validation, and error handling mechanisms within data pipeline.

Optimize Data Integration Processes:

· Enhance the performance, optimization of data workflows to meet SLAs, scalability of data integration processes and cost-efficiency on AWS cloud infrastructure.

· Identify and resolve performance bottlenecks, fine-tuning queries, and optimizing data processing to enhance Redshift's performance

· Regularly review and refine integration processes to improve efficiency.

· Support Business Intelligence and Analytics:

· Translate business requirements to technical specifications and coded data pipelines

· Ensure timely availability of integrated data for business intelligence and analytics.

· Collaborate with data analysts and business stakeholders to meet their data requirements.

Maintain Documentation and Compliance:

· Document all data integration processes, workflows, and technical & system specifications.

· Ensure compliance with data governance policies, industry standards, and regulatory requirements.

WHAT WILL THIS PERSON BE WORKING ON

· 5+ years of experience in data engineering, database design, ETL processes, and data warehousing.

· 3+ years of experience with AWS tools and technologies (S3, EMR, Glue, Athena, RedShift, RDS, Spectrum and Airflow)

· 2+ years of experience with CI/CD tools.

· Strong knowledge of data storage and processing technologies, including databases and data lakes based distributed computing frameworks (e.g., Hadoop, Spark).

· 3+ in programming languages such as Python, Java, or Scala.

· Nice to have Informatica Cloud tool experience (IDMC)

· Nice to have Agentic AI experience with Amazon Kiro

WANTS

Primary Skills : AWS EMR , GLUE, AIRFLOW, ICEBERG, REDSHIFT, RDS , IDMC, AMAZON KIRO and CI/CD

Candidate should Design and Develop Data Pipelines and Support them.

Candidate will work from Torrance CA location, 4 days onsite and 1 day Work from Home, Should be able to come work all five days if required.

Complete Job Requirement:

Experience:

· 5+ years of experience in data engineering, database design, ETL processes, and data warehousing.

· 3+ years of experience with AWS tools and technologies (S3, EMR, Glue, Athena, RedShift, RDS, Spectrum and Airflow)

· 2+ years of experience with CI/CD tools.

· Strong knowledge of data storage and processing technologies, including databases and data lakes based distributed computing frameworks (e.g., Hadoop, Spark).

· 3+ in programming languages such as Python, Java, or Scala.

· Nice to have Informatica Cloud tool experience (IDMC)

· Nice to have Agentic AI experience with Amazon Kiro

Detailed Responsibilities:

· Develop and Maintain Data Integration Solutions:

· Design and implement data integration workflows using AWS Glue/EMR,AWS MWAA(Airflow), Lambda, Redshift

· Demonstrate proficiency in Pyspark, Apache Spark and Python for data processing large datasets

· Ensure data is accurately and efficiently extracted, transformed, and loaded into target systems.

Similar jobs