Data Engineer (Python / PySpark) - only w2
Quick Overview
Job Description
Role Summary:
Financial Crimes Technology team is evolving toward more in-house build capabilities and reducing dependency on legacy vendor/tooling approaches.
This role will support data engineering and platform development for financial crimes use cases (AML, investigations, sanctions, fraud, KYC) by building scalable pipelines, improving data quality, and enabling analytics/reporting and downstream applications.
Key Responsibilities:
Build and maintain batch and/or streaming data pipelines supporting financial crimes initiatives.
Develop data transformations using Python + PySpark and optimize performance for large datasets.
Apply strong understanding of Apache Spark architecture (executors, partitions, shuffles, joins, caching) to improve performance
Partner with business and technical stakeholders to translate requirements into data models, mappings, and curated datasets.
Support ingestion from multiple sources (transactional systems, case management, reference data, etc.).
Implement data quality checks, reconciliation, and controls to ensure auditability and reliability.
Contribute to modernization efforts (legacy → in-house build) including migration planning and redesign.
Create documentation for pipelines, logic, and operational runbooks.
Work within Agile delivery (Jira), supporting sprint execution and delivery timelines.
Required Skills:
5+ years of experience in data engineering / ETL / data platform development
Strong hands-on development in:
Python
PySpark / Apache Spark
Advanced SQL
Experience working with large-scale data sets and performance tuning.
Strong understanding of data concepts: data modeling, lineage, metadata, governance
Experience supporting regulated environments with emphasis on controls and audit readiness
Strong communication skills (ability to work with both engineering + business partners)
Preferred Skills (Nice to Have):
Experience running PySpark workloads on Google Cloud Platform (Google Cloud Platform)
Dataproc, BigQuery, Google Cloud Storage (GCS), etc.
Experience with cloud native Big Data platforms
Knowledge of data governance, security, and compliance practices
Experience with CI/CD pipelines for data engineering workloads
Orchestration: Airflow (or similar scheduling tools)
Streaming: Kafka
Lakehouse/Warehouse: Databricks / Snowflake / BigQuery
CI/CD + DevOps: Git, pipelines, automation, release management
Data governance/security: encryption, access controls, data masking, PII handling
Prior Financial Crimes domain: AML / sanctions / fraud / investigations / KYC
Domain Experience (Highly Valued):
Experience supporting AML, Transaction Monitoring, Investigations, Sanctions screening, Fraud, or similar risk/compliance functions.
Familiarity with regulatory expectations and strong documentation discipline.
Skills
Similar jobs
ETL Data Engineer with Security Clearance
Noblis · Reston, United States
1 hour ago$90.7k - $141.8k/yrData Engineer Snowflake
Codeforce 360 · United States
1 hour agoUrgent Requirement – Data Lead / Senior Data Engineer
METANLYTICS LLC · Chicago, United States
1 hour agoData Engineer
Archon Resources · Tulsa, United States
1 hour agoData Engineer with Security Clearance
Expression Networks, LLC · Arlington, United States
1 hour ago$10k/yrQNXT Data Engineer
Stellent IT LLC · United States
1 hour ago