Quick Overview
Job Description
Key Technical Responsibilities
Design, develop, and maintain data pipelines and ETL/ELT workflows supporting machine learning, NLP, retrieval, and advanced analytics. Prepare, transform, curate, and validate structured and unstructured grants and financial data for AI/ML applications. Perform feature engineering, document processing, chunking, embeddings generation, and dataset preparation for AI-assisted analytics and retrieval workflows. Develop reproducible data transformations and workflows using Python, SQL, and applicable data engineering technologies.
Implement and maintain data quality, validation, provenance, lineage, versioning, metadata, and source traceability throughout the data lifecycle. Prepare and integrate data for search, retrieval-augmented generation (RAG), visualization, analytics, and AI-assisted decision support. Troubleshoot data pipeline, transformation, integration, and data-quality issues across development and testing environments. Collaborate with multidisciplinary data engineering, data science, software engineering, and AI/ML teams during iterative development and testing.
Develop technical documentation covering data architecture, pipelines, transformations, schemas, features, embeddings, dependencies, and operational procedures. Support deployment and operation of data engineering capabilities within secure Federal, on-premises, cloud, or hybrid environments. Ensure data engineering solutions comply with Federal security, privacy, accessibility, records-management, data-ownership, and governance requirements.
Required Technical Qualifications
Demonstrated experience building, integrating, and operating data pipelines for machine learning, NLP, information retrieval, or advanced analytics. Hands-on experience with feature engineering, document processing, embeddings, curated datasets, data transformation, and reproducible data workflows. Strong proficiency in Python and SQL for data engineering, data transformation, automation, and model-support workflows. Experience working with structured and unstructured data and preparing data for AI/ML or analytics applications.
Experience implementing data quality controls, data validation, provenance, metadata, versioning, lineage, and source traceability for AI/ML or data-intensive systems. Experience troubleshooting and optimizing data pipelines, integrations, transformations, and data-processing workflows. Ability to develop clear technical documentation covering data pipelines, schemas, transformations, dependencies, and operational processes. Strong technical communication, problem-solving, and cross-functional collaboration skills.
Similar jobs
- AG
Senior Data Scientist
NewACI Group, Inc.
United States🇺🇸Remote23 hours agoDynamoDBAPI GatewayAWS+18Technology - RT
Data Scientist II
NewRussell, Tobin & Associates
Mountain View, CA🇺🇸$65 - $72/hrHybrid23 hours agoSQLGitPythonTechnology - SP
Data Scientist with Security Clearance
NewSPA
Washington, DC🇺🇸$90k - $105k/yrHybrid23 hours agoMachine LearningTechnology - PC
Financial Data Scientist - Capital Markets
NewPalmetto Clean Technology
New York🇺🇸Hybrid7 hours agoSQLNumPyContinuous Improvement+8Technology - KT
Principal Data Scientist
NewKforce Technology Staffing
Juno Beach, FL🇺🇸Hybrid23 hours agoSQLAWSMachine Learning+10Technology - B&
Senior Exploitation Specialist / Data Scientist with Security Clearance
NewBart & Associates
Sneads Ferry, NC🇺🇸Hybrid23 hours agoPythonTechnology