Why This Role Stands Out
This hybrid Data Engineer role offers a fantastic opportunity to build and maintain critical data pipelines in the exciting media and video space, supporting cutting-edge AI research. You'll thrive here if you're a skilled Python and SQL developer with a passion for data quality and automation, eager to contribute to impactful projects within a collaborative environment. Apply now to leverage your expertise and grow your career!
Quick Overview
Job Description
Job Description
- The main function of the Data Engineer is to develop, evaluate, test, and maintain architectures and data solutions within our organization. The typical Data Engineer executes plans, policies, and practices that control, protect, deliver, and enhance the value of the organization s data assets.
Day-to-day responsibilities
- Build and maintain data mitigation pipelines that scan, classify, and remediate research datasets before research use.
- Create and manage Hive tables and namespaces for mitigated outputs.
- Move data across Manifold, S3, and NFS
- Debug and triage pipeline failures (workflow conflicts, rate limits, cross-namespace query errors) and document runbooks.
- Request, track, and validate dataset read/write access grants for mitigation workflows.
- Author and iterate on agentic skills/automation (e.g., fdp-mitigate) to make mitigation repeatable and self-service.
- Partner with AI researchers and XFN reviewers on dataset readiness, and report status to the data engineering lead
Key Projects: AI dataset mitigations (e.g. CSAM image/video mitigation. copyright) using agentic tooling. MAANG Experience Highly Wanting.
Must-Have Skills
- Strong Python, SQL/Presto for large-scale batch data processing.
- hands-on data pipeline and workflow engineering orchestration, restart ability, idempotency, debugging at scale.
- practical data-quality and validation rigor (schema checks, dedupe, threshold gating) with clear documentation habits.
Nice-to-Have Skills
- Media/video data processing decode, clipping, frame extraction, embeddings.
- Experience with AI coding agents / prompt-and-skill authoring to automate repetitive data workflows.
Years of Experience: 04 - 05 Years
Education: BS in Computer Science, Data Engineering, or a related technical field; equivalent hands-on data pipeline experience accepted in lieu of a degree. Advanced degree not required.
Interviews: 02 Rounds | Technical / Behavioral | 30 - 45 Minutes.
Similar jobs
- AI
Data Engineer, Infrastructure FinOps with Security Clearance
NewAnduril Industries
Costa Mesa, CA🇺🇸$146k/yrHybridYesterdayDockerGCPSQL+18Technology - JS
Elastic Data Engineer with Security Clearance
NewJCS Solutions LLC
Southern Md Facility, MD🇺🇸$135k - $162k/yrHybridYesterdayDockerSQLAWS+11Technology - TA
Project Engineer - Data Analysis
NewAuto ApplyTrue Anomaly
Denver🇺🇸Hybrid11 hours agoSQLSQL ServerAzure+9Engineering - P-
Super Day - Data Engineering Internship
NewAuto ApplyPerpay - Career's Page
Philadelphia🇺🇸Hybrid14 hours agoSQLAWSAirflow+5 - AI
Senior Data Engineer (Seattle)
NewAuto ApplyAirwallex
US - Seattle🇺🇸Hybrid14 hours agoMySQLOracleSQL+9Technology - JA
Senior Data Systems Engineer
NewAuto ApplyJanuary
New York City🇺🇸Hybrid12 hours agoData PipelineLESSTechnology - BL
Senior Data Engineer, Product
NewAuto ApplyBlock
Bay Area🇺🇸Hybrid10 hours agoSQLETLMachine Learning+8Technology - AX
Principal Data Engineer
NewAuto ApplyAxonius
Remote US🇺🇸Remote13 hours agoSQLAWSETL+8Technology - AI
Data Engineer, Infrastructure FinOps
NewAuto ApplyAnduril Industries
Costa Mesa🇺🇸Hybrid14 hours agoDockerGCPSQL+24Technology - AL
Senior Knowledge Engineer, Ontology and Data Standards (Contract)
NewAuto ApplyAltos Labs
San Francisco Bay Area🇺🇸Hybrid15 hours agoExpressAWSGit+6Engineering - 66
Associate Data Engineer, Gradient Specialist
NewAuto Apply66degrees
Chicago🇺🇸Hybrid13 hours agoSpringScrumGoogle Cloud+2Technology - JO
Data Engineer - Finance Systems & Analytics
NewJooble
United States🇺🇸$90 - $95/hrRemote2 days agoSQLETLPower BI+1Technology