Quick Overview
Job Description
Job Title: AWS Lakehouse Data Engineer
Work Model: Remote – offsite
Responsibilities
- Build and operate data pipelines (batch and streaming) from APIs, relational databases, file drops, event streams, and external partners.
- Design, implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce analytics-ready datasets for reporting, visualization, and machine learning.
- Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
- Improve pipeline reliability through automated testing, orchestration, monitoring, retries, and operational runbooks.
- Design and implement a Delta Lakehouse-style data platform on AWS using native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
- Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
- Implement SQL-like table reliability features including ACID transactions, schema evolution, snapshot isolation, and time travel using Apache Iceberg.
- Enable fast, interactive queries of lakehouse data via AWS-native services like Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
- Optimize performance and cost through partitioning, file sizing, caching, lifecycle policies, and separating compute from storage.
- Establish standardized environments for development, testing, and production with consistent configuration and controlled promotion.
- Implement data governance, access control, lineage, and quality measures utilizing AWS-native services including AWS Lake Formation, AWS Glue Data Catalog, IAM, KMS.
- Create a metadata repository with cataloging, ownership, classification, tagging, and discoverability features.
- Enable end-to-end data lineage for audit and regulatory compliance.
- Apply policy-based access, least privilege, data classification, retention, encryption, and secure handling controls.
- Build data quality checks for freshness, completeness, validity, and anomaly detection, and publish SLA/SLO metrics.
- Automate AWS provisioning with Infrastructure as Code (IaC), develop CI/CD pipelines for data components, and ensure platform observability.
- Work collaboratively with cross-functional teams and maintain high-quality engineering documentation.
Requirements
- Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or related field, or four (4) years of equivalent practical experience.
- Six (6) years of relevant experience.
- Hands-on experience building AWS-native data lake or lakehouse architectures on Amazon S3.
- Strong experience developing production ETL/ELT pipelines with Python and PySpark, including data modeling, transformation, and performance tuning.
- Hands-on experience with Apache Iceberg, including ACID transactions, schema evolution, time travel, and query optimization.
- Advanced SQL skills supporting analytical workloads, reporting, and data visualization.
- Proven experience with data governance, cataloging, lineage, and access control using AWS services.
- Knowledge of AWS security fundamentals: IAM, KMS, secrets management, network security, logging, SDLC.
- Proven experience with Infrastructure as Code (IaC) and operating data platforms across environments.
- Experience with CI/CD pipelines for data workflows with testing, deployment, environment promotion, and rollback.
- Troubleshooting distributed data workloads, performance optimization, and cost management skills.
- Excellent collaboration and communication skills to coordinate with cross-team stakeholders.
Would Be Nice to Have
- Experience with Databricks, Delta Lake, migrating workloads to AWS-native services, and Apache Iceberg.
- Familiarity with AWS Step Functions, MWAA, Kinesis, DMS, Lambda, MSK, or similar services.
- Experience with modern DevOps tools: Git, Terraform, CloudFormation, Jenkins, CodePipeline, GitHub Actions, Docker.
- Knowledge of BI and visualization tools like Amazon QuickSight, Tableau, Power BI.
- Familiarity with AI-assisted coding tools such as GitHub Copilot, ChatGPT, Cursor, or Kiro.
- Knowledge of graph modeling, ontology, taxonomy, entity resolution, and hybrid retrieval techniques.
System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.
System One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.
#LI-KA1
#M1
Ref: #851-Rockville-S1
Similar jobs
- UI
AI Ready Data Engineer
NewUnikon IT
Seattle, WA🇺🇸Hybrid20 hours agoMachine LearningApacheApache Spark+1Technology - DT
Data Engineer (Oracle apps) (on W2 only)
NewDigipulse Technologies, Inc
Durham, NC🇺🇸Hybrid20 hours agoOracleSQLAWS+7Technology - TE
Data Engineer
NewTECHProjects
United States🇺🇸Remote20 hours agoSQLSQL ServerETL+4Technology - GR
Geospatial Data Manager with Security Clearance
NewGRVTY
Springfield, VA🇺🇸Hybrid20 hours agoOracleSQLTableau+1Technology - HR
Senior Data Engineer (Pyspark)
NewHRConnects
Jersey City, NJ🇺🇸On-site20 hours agoDynamoDBSQLAWS+1Technology - LT
Snowflake Data Engineer with Cortex
NewLTM
United States🇺🇸Remote20 hours agoSQLAWSETL+3Technology