Haystack
← Back to Jobs
Technology

Senior Data Scientist

Get A WhizWoodlawn, MD🇺🇸United StatesPosted 17 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Senior Data Scientist

Introduction:

We are looking for a Senior Data Scientist with strong experience in Python, SQL, Natural Language Processing (NLP), text processing, and data matching/entity resolution. The ideal candidate will have hands-on experience building data processing pipelines, cleansing large datasets, and developing advanced data matching solutions.

Responsibilities:

  • Develop and maintain data processing, analytics, and entity resolution pipelines using Python and SQL.
  • Clean, transform, and manage large datasets from multiple complex sources.
  • Develop solutions for NLP, text processing, information extraction, and data matching.
  • Optimize complex SQL queries and database operations for performance and scalability.
  • Perform data validation, testing, deployment, and post-production monitoring.
  • Participate in code reviews and follow version control, data security, and coding best practices.
  • Work independently and collaborate with technical and business teams to deliver data solutions.

Requirements:

Required Skills:

  • 10+ years of overall IT experience.
  • Strong hands-on experience with Python and SQL.
  • Experience with Natural Language Processing (NLP), Text Processing, and Information Extraction.
  • Experience with Named Entity Recognition (NER) and address standardization.
  • Strong understanding of data matching / entity resolution techniques.
  • Experience with Blocking and Indexing, String Distance Metrics, Phonetic Encoding, TF-IDF and Cosine Similarity, and Regex / Pattern Matching.
  • Experience with spaCy and Scikit-Learn.
  • Experience with entity resolution/linkage libraries such as Splink, FastLink, Dedupe, or RecordLinkage.
  • Strong SQL query optimization and database skills.
  • Experience with Git/version control, code reviews, data security, and reproducible development practices.

Preferred / Nice to Have:

  • Experience working on Federal, State, or Local Government data projects.
  • Experience with PostgreSQL, DB2, Oracle, SQL Server, Hadoop, and flat-file data.
  • Experience with Jenkins and CI/CD for data pipelines.
  • Experience with pipeline automation and job scheduling/monitoring tools.
  • Experience explaining complex data matching algorithms and business rules to non-technical stakeholders.
  • Strong analytical, problem-solving, communication, and collaboration skills.

Skills

Oracle
SQL
SQL Server
NLP
Scikit-learn
Git
Hadoop
Jenkins
PostgreSQL
Python

Similar jobs