Quick Overview
Job Description
Data Engineer
Remote
About the Role
We are seeking a Data Engineer / Data Operations Analyst to join our Onshore team, responsible for managing, monitoring, and operating data pipelines that support US healthcare quality reporting and analytics. This role requires close, day-to-day collaboration with an Offshore engineering team, hands-on experience running and troubleshooting data processing engines, and solid working knowledge of AWS file transfer and storage concepts (S3). Candidates should bring or be able to quickly pick up domain understanding of US healthcare quality programs (e.g., Star Ratings, HEDIS, quality measures, rebate/bonus calculations).
Key Responsibilities
Serve as the primary Onshore point of contact, coordinating daily work, priorities, and handoffs with the Offshore data engineering/operations team, ensuring smooth 24-hour work continuity across time zones.
Understand, operate, and monitor existing data processing engines/pipelines, including scheduled and on-demand job runs, ensuring successful completion and timely resolution of failures.
Run, restart, and troubleshoot data pipeline jobs (batch/streaming) in production and non-production environments; escalate and document recurring issues.
Manage and monitor AWS-based file transfers, including inbound/outbound file exchanges with external partners .
Work with Amazon S3 concepts - bucket structure, object versioning, lifecycle policies, access permissions (IAM/bucket policies), and data partitioning - to ensure secure and reliable data storage and retrieval.
Perform data validation, reconciliation, and quality checks on incoming/outgoing files to ensure completeness and accuracy.
Support and maintain ETL/ELT workflows feeding into downstream reporting, analytics, and compliance systems.
Document pipeline runbooks, SOPs, and operational procedures for repeatable, auditable processes.
Create and maintain technical documentation, including process flows, data flow diagrams, pipeline architecture notes, and knowledge-transfer materials for both Onshore and Offshore teams.
Design and build new operational processes/workflows from scratch where none exist - identifying gaps, defining steps, and standardizing repeatable procedures for data ingestion, transformation, and file handling.
Collaborate with business/quality analysts to understand data requirements tied to US healthcare quality metrics (e.g., Star Ratings, HEDIS measures, CMS reporting cycles).
Participate in incident management, root cause analysis, and continuous improvement initiatives for data operations.
Support on-call rotations or extended coverage windows as needed to align with Offshore team handoffs
Required Qualifications
Bachelor's degree in Computer Science, Information Systems, Data Engineering, or related field (or equivalent practical experience).
5+ years of experience in data engineering, data operations, or production data pipeline support.
Proven experience collaborating with Offshore/distributed teams, including clear communication, task handoff, and status reporting across time zones.
Hands-on experience operating and troubleshooting data engines/pipelines (e.g., Spark, Airflow, Informatica, Talend, or similar ETL/orchestration tools).
Familiarity with US healthcare industry data and quality reporting concepts - Medicare Advantage, HEDIS, quality/rebate measures, or similar NCQA related programs strongly preferred.
Strong troubleshooting mindset with ability to triage and resolve production data issues under time pressure.
Demonstrated ability to build new processes/workflows from the ground up and create clear technical documentation (process guides, SOPs, runbooks) for recurring operational tasks.
Excellent written and verbal communication skills, with ability to translate technical issues for both Onshore and Offshore stakeholders.
Strong Problem-solving skill.
Technical Skills
Programming/Scripting: Python, PySpark, Bash scripting, AWK command
Databases: MySQL, Amazon Redshift, and general relational database concepts (SQL, query optimization, schema design)
Cloud & Storage: AWS (EBS,S3, file transfer mechanisms, IAM basics)
Version Control: Git (branching, merging, pull requests, code review workflows)
Understanding of Containerization: Docker (building, running, and managing containers for data pipeline components)
Big Data Processing: PySpark for distributed data processing and transformation at scale
Preferred Qualifications
Experience with cloud-based data warehousing (Redshift).
Exposure to healthcare compliance and data privacy standards (HIPAA).
Experience with job scheduling/monitoring tools (Control-M, Airflow, AWS Step Functions, etc.).
Prior experience in a healthcare payer, health plan, or Medicare Advantage environment.
Familiarity with CI/CD practices for data pipeline deployment.
Knowledge of FHIR structures, CCDA etc., will be a plus
Optionally knowledge of Databricks is an added advantage
Skills
Similar jobs
Databricks Data Engineer
Microgreen Technologies LLC · United States
8 minutes agoSnowflake Data Engineer
CLPS Global · New York, United States
8 minutes agoData Engineer with Security Clearance
Base-2 Solutions, LLC · Miami, United States
31 minutes ago$10k/yrData Engineer – Data & AI, Supply Chain
InfoVision, Inc. · United States
31 minutes agoSenior Integration & Data Engineer (Developer III)
INSPYR Solutions · Charlotte, United States
31 minutes ago$45 - $50/hrData Engineer
DCI Solutions · Washington, United States
32 minutes ago