Haystack
← Back to Jobs
Technology
IT

Senior Data Scientist / Machine Learning Engineer (Document Intelligence & OCR)

isolve technology incUnited States🇺🇸United StatesPosted Sep 30, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
DynamoDBSQLAWSMLOpsMachine LearningNLPPostgreSQLPython

Job Description

Senior Data Scientist / Machine Learning Engineer (Document Intelligence & OCR)

Position Purpose

The Senior Data Scientist / Machine Learning Engineer will lead the design, development, and deployment of end-to-end Document Intelligence solutions for the U.S. Fish and Wildlife Service (USFWS). In this role, you will build robust document-ingestion, OCR, field-extraction, free-text remediation, and classification pipelines to process permits, certificates, and legacy records. A core focus will be implementing confidence-based human-review workflows and maintaining strict, precise source traceability across all extracted data points.

Key Responsibilities

  • Pipeline & OCR Engineering: Build and optimize end-to-end document processing pipelines, combining OCR, layout-aware models, and NLP techniques to extract structured fields and remediate unstructured free text from permits, certificates, and scanned legacy records.

  • Traceability & Bounding Box Mapping: Establish precise document-coordinate mapping (bounding boxes/character-level offsets) to map extracted fields back to original source documents for complete auditability.

  • Human-in-the-Loop Workflow Design: Design confidence-scoring mechanisms and route low-confidence or high-risk extractions to human-in-the-loop (HITL) review interfaces.

  • Model Evaluation & Threshold Tuning: Perform rigourous evaluation on labeled datasets, analyzing precision, recall, and edge cases to calibrate operational confidence thresholds.

  • Cloud Architecture & Data Modeling: Implement scalable processing pipelines using AWS cloud services, PostgreSQL/SQL databases, document data models, and RESTful APIs.

  • Security & Compliance: Ensure all solutions comply with federal security frameworks (FISMA, FedRAMP, NIST SP 800-53) and Privacy Act controls, designing explainable and auditable automated decisions.

Required Capabilities

  • Data Science & ML Expertise: Senior-level experience in applied machine learning/data science focused on OCR, document understanding, information extraction, NLP, or text classification.

  • Python & Evaluation Mastery: Advanced Python skills with demonstrated experience in model evaluation, precision/recall metrics, threshold calibration, and structured error analysis.

  • Traceable HITL Systems: Proven track record of building traceable, auditable human-review workflows for low-confidence AI predictions.

  • Data Systems & Integration: Strong working knowledge of SQL/PostgreSQL, semi-structured document models, API development, and cloud object storage (e.g., S3).

  • Governance & Security: Experience designing audit trails for automated decision systems and safeguarding sensitive or PII data.

Preferred Qualifications

  • Education: Bachelor’s degree or higher in Computer Science, Data Science, Machine Learning, Computational Linguistics, or a related quantitative field.

  • AWS Ecosystem: Hands-on experience with AWS AI/ML and serverless services, such as Textract, Bedrock, SageMaker, Lambda, S3, and DynamoDB (or federal cloud equivalents).

  • Government Document Processing: Prior experience extracting structured data from government forms, permits, environmental certificates, scientific publications, or historical scanned archive collections.

  • Advanced Document Understanding: Experience with layout-aware vision-language models, image preprocessing (deskewing, binarization, noise reduction), taxonomy alignment, entity linking, and MLOps practices.

  • Federal Compliance & Controls: Familiarity with FISMA, NIST SP 800-53, FedRAMP compliance, the Privacy Act, model explainability, and AI model risk management principles.

Similar jobs