Haystack
← Back to Jobs
Technology

Senior Data Engineer

Smart SynergiesBethesda, MD🇺🇸United StatesPosted 26 Jul 2026

Why This Role Stands Out

This hybrid Senior Data Engineer role offers significant impact by designing and implementing scalable data pipelines within a cutting-edge AWS Lakehouse architecture, empowering data-driven decisions across the organization. You'll thrive here if you're a proactive engineer eager to advance your skills in cloud data platforms and collaborate with a dynamic team. Apply today to contribute to a company renowned for its innovative technology solutions!

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Role Summary
We are looking for a driven and hands-on Data Engineer to design, implement, and support scalable, production-level data pipelines within our AWS-based data ecosystem. This position is responsible for full lifecycle pipeline delivery—from ingesting data from enterprise systems to producing analytics-ready datasets—while collaborating closely with Data Architects, Analytics teams, and business partners. The role is critical in ensuring reliable, high-quality data is available across the organization to power analytics and informed decision-making. As a member of the core data platform team, you will also contribute to advancing our AWS Lakehouse architecture and establishing engineering best practices.

Key Responsibilities

  • Pipeline Engineering: Develop and maintain scalable ELT pipelines using reusable ingestion frameworks that support both batch and event-driven processing across multiple data sources such as ERPs, APIs, vendor feeds, and relational systems.

  • AWS Data Platform Development: Design and enhance AWS-native data solutions aligned with a Medallion (Bronze/Silver/Gold) Lakehouse architecture, with an emphasis on performance optimization and cost management.

  • Infrastructure & DevOps: Build and manage AWS infrastructure using Terraform, and support CI/CD processes through GitOps methodologies, including automated testing, monitoring, alerting, and system recovery capabilities.

  • Data Modeling & Transformation: Create and maintain dimensional models and Gold-layer datasets using SQL, Python, and PySpark. Implement scalable ingestion processes with strong handling of schema evolution, auditing, and performance tuning techniques such as partitioning, clustering, and materialization.

  • Data Quality & Governance: Integrate automated data validation, anomaly detection, and lineage tracking into pipelines, while contributing to metadata management practices.

  • Reporting & BI Enablement: Diagnose and resolve complex data issues, ensuring pipelines deliver data optimized for Power BI. Work closely with BI developers to align data models with reporting needs and troubleshoot dashboard-related challenges.

  • Cross-Functional Collaboration: Partner with architecture and analytics teams to translate business requirements into technical solutions, and actively participate in code reviews, sprint planning, and architectural discussions.

Qualifications

  • Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related discipline, or equivalent practical experience.

  • At least 7 years of experience in data engineering, with strong hands-on expertise in AWS services and distributed data technologies.

  • Advanced proficiency in Python and SQL, including experience with Spark/PySpark.

  • Proven experience building and maintaining AWS-based data pipelines using services such as Glue, Step Functions, Lambda, S3, Athena, SNS, SQS, and Redshift.

  • Experience with event-driven data architectures.

  • Hands-on experience processing and managing large-scale vendor data feeds.

  • Practical knowledge of Medallion architecture within a data lake or Lakehouse environment.

  • Experience using Terraform for infrastructure-as-code deployments.

  • Familiarity with CI/CD tools such as Bitbucket, GitHub, or AWS CodePipeline for pipeline deployment.

  • Strong understanding of data warehousing principles, including star schema design, dimensional modeling, and slowly changing dimensions (SCD).

  • Experience integrating data with Power BI or comparable business intelligence tools.

  • Working knowledge of UNIX/Linux environments, including shell scripting.

  • Experience supporting production systems, including monitoring, troubleshooting, and on-call responsibilities.

  • Familiarity with Agile development methodologies.

Preferred Qualifications

  • Experience integrating data from Oracle EBS.

  • Familiarity with data quality tools such as Great Expectations or dbt testing frameworks.

  • Experience with AWS CDK or CloudFormation alongside Terraform.

  • Knowledge of data cataloging and lineage tools such as Alation or Collibra.

  • AWS certifications (e.g., AWS Certified Data Engineer, AWS Certified Solutions Architect).

Skills

Oracle
SQL
Shell
AWS
Agile
CDK
CloudFormation
Power BI
Python
Redshift
Terraform
dbt

Similar jobs