Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Baltimore, MD, United States
Posted
19 hours ago
SQLScalaAWSETLEncryptionSnowflakeAirflowApacheApache SparkCloudFormationData PipelineHIPAAJavaPythonRedshiftTerraform
Job Description
Role: Data Engineer
Location: Baltimore City, MD (Hybrid)
Duration: 12+ months project with 1 one-year renewal option
Job Description:
Hands-On Data Pipeline Development
- Design, code, and deploy ETL/ELT pipelines across bronze, silver, and gold layers of the Data Lakehouse.
- Build ingestion pipelines for structured (SQL), semi-structured (JSON, XML), and unstructured data using PySpark/Python programming language using AWS Glue or EMR.
- Implement incremental loads, deduplication, error handling, and data validation.
- Actively troubleshoot, debug, and optimize pipelines for scalability and cost efficiency.
EDW & Data Lake Implementation
- Develop dimensional data models (Star Schema, Snowflake Schema) for analytics and reporting.
- Build and maintain tables in Iceberg, Delta Lake, or equivalent OTF formats.
- Optimize partitioning, indexing, and metadata for fast query performance.
Healthcare Data Integration
- Build ingestion and transformation pipelines for EDI X12 transactions (837, 835, 278, etc.).
- Implement mapping and transformation of EDI data with FHIR and HL7 frameworks.
- Work hands-on with AWS Health Lake (or equivalent) to store and query healthcare data.
Data Quality, Security & Compliance
- Develop automated validation scripts to enforce data quality and integrity.
- Implement IAM roles, encryption, and auditing to meet HIPAA and CMS compliance standards.
- Maintain lineage and governance documentation for all pipelines.
Collaboration & Delivery
- Work closely with the Lead Data Engineer, analysts, and data scientists to deliver pipelines that support enterprise-wide analytics.
- Actively contribute to CI/CD pipelines, Infrastructure-as-Code (IaC), and automation.
- Continuously improve pipelines and adopt new technologies where appropriate.
Minimum Qualification:
Specialized experience: The candidate should have experience as data engineer or similar role with a strong understanding of data architecture and ETL processes. The candidate should be proficient in programming languages for data processing and knowledgeable of distributed computing and parallel processing.
- 3+ years hands-on experience in building, deploying, and maintaining data pipelines on AWS or equivalent cloud platforms.
- Strong coding skills in Python and SQL (Scala or Java a plus).
- Proven experience with Apache Spark (PySpark) for large-scale processing.
- Hands-on experience with AWS Glue, S3, Redshift, Athena, EMR, Lake Formation.
- Strong debugging and performance optimization skills in distributed systems.
- Hands-on experience with Iceberg, Delta Lake, or other OTF table formats.
- Experience with Airflow or other pipeline orchestration frameworks.
- Practical experience in CI/CD and Infrastructure-as-Code (Terraform, CloudFormation).
- Practical experience with EDI X12, HL7, or FHIR data formats.
- Strong understanding of Medallion Architecture for data lake houses.
- Hands-on experience building dimensional models and data warehouses.
- Working knowledge of HIPAA and CMS interoperability requirements.
Similar jobs
- AA
Data Engineer/Sr Data Engineer, IT Analytics
American Airlines
Fort Worth, TX🇺🇸Hybrid1 week agoSQLMachine LearningAgile+3Technology - MM
Data Engineer
NewMitchell Martin, Inc.
Rosemont, IL🇺🇸$87 - $97/hrOn-site19 hours agoTechnology - CG
Senior Data Engineer
NewCharter Global, Inc.
Dallas, TX🇺🇸Hybrid19 hours agoSQLAWSETL+10Technology - HS
Senior Data Engineer
Hierarch Soft Technologies, Inc.
New York, NY🇺🇸Hybrid3 weeks agoPL/SQLSQLETL+6Technology - SS
Data Engineer || United States citizenship is required
NewSolveIT Services Inc
United States🇺🇸Hybrid19 hours agoSQLAWSAzure+1Technology - SA
Data Engineer - W2 & C2C - New York - Direct Client
NewSANS
New York, NY🇺🇸Hybrid19 hours agoSQLAzureJira+3Technology