Why This Role Stands Out
You'll have a significant impact by building robust data pipelines that power cutting-edge AI legal software, offering excellent opportunities for professional growth in a remote setting. This role is ideal for a self-driven Data Engineer with a passion for tackling complex data challenges and a strong background in Python, SQL, and AI-driven data extraction. Apply to join a reputable company and contribute to transforming legal research.
Quick Overview
Job Description
About the company
Our client builds AI-powered legal software for research and litigation workflows. Its products use large collections of court records and legal documents to help turn complex information into structured insights.
The role / why it matters
You will own the data pipelines that ingest, normalize, and enrich caselaw and docket records. Legal source data is messy, often arriving as scanned filings or semi-structured documents. The reliable signal you create will support downstream research and AI products.
What you'll do
• Build and maintain production pipelines for PACER, NYSCEF, state court systems, published opinions, and other legal records.
• Turn PDFs, scanned filings, XML, and HTML into clean, queryable data.
• Create LLM-assisted extraction workflows for legal text and validate results against source documents.
• Run statistical analyses, improve data quality, and optimize SQL queries.
• Support RAG and semantic search workflows in partnership with full-stack engineers.
• Operate and improve AWS S3 and RDS data systems within established security and compliance guardrails.
What we're looking for
• 4-8 years of experience building and operating production data pipelines with large, dirty datasets, and quickly finding signal in them.
• Strong Python and SQL skills for pipeline engineering and data analysis.
• Hands-on experience with legal data such as PACER records, court dockets, or court filings.
• Experience with AI for data, including RAG, semantic search, and LLM-assisted extraction, plus PDF parsing, OCR, and document extraction tooling.
• A bachelor's degree in Computer Science, Data Science, or a related field.
• Works fully autonomously and thrives with minimal management direction; stays curious about messy data problems and critically reviews AI-assisted code and outputs.
Bonus points
• Experience with regulatory or government data sources.
• Experience with vector databases or semantic-search tooling such as Pinecone or Voyage AI.
Compensation and benefits
• Base salary: $180,000-$220,000 USD.
• Equity: 0.15%-0.25%.
Location / work model
• Remote within the United States, with at least four hours of overlap with Eastern Time each workday.
• Strong preference for candidates based in the New York City area.
• Plan for quarterly travel to New York City.
• Visa transfers may be considered. New visa sponsorship is not available.