Haystack
← Back to Jobs
Technology

Senior Data Scientist

International Software Systems, IncWoodlawn, MD🇺🇸United StatesPosted 14 Aug 2026

Quick Overview

Work Type
On Site
Level
Mid Senior

Job Description

Key Required Skills:

  • Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction concepts including Named Entity Recognition, Blocking and Indexing, String Distance Metrics, TF-IDF/Cosine Similarity, Phonetics Encoding, Address Standardization

  • Solid Python, Regex, and SQL experience

  • Excellent Communication skillsPosition Description:

    • Develop Analytics Solutions: Design, implement, and maintain advanced data processing and entity resolution pipelines using Python and SQL on enterprise data platforms.

    • Data Hygiene & Management: Clean, transform, and manage large-scale datasets from diverse, complex sources, ensuring absolute data integrity, reliability, and security.

    • Performance Optimization: Optimize complex SQL queries and database operations to ensure efficient data access, processing, and scalability.

    • Engineering Standards: Actively participate in code reviews, enforce version control, and uphold best practices for code quality, reproducibility, and data privacy.

    • End-to-End Delivery: Support data validation, testing, deployment, and post-implementation monitoring in a fast-paced environment.

    Skills Requirements:

    Foundation for Success (Basic Qualifications)

    • Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, or Information Science with experience in NLP, Text Processing, Information Extraction, Python, SQL, Regex and specialized libraries/frameworks.

    • Overall 10+ years’ experience in IT industry

    Factors To Help You Shine (Required Skills)

    **Selected candidate must be able to obtain and maintain a public trust clearance**
    **Selected candidate must be willing to work on-site in Woodlawn, MD 5 days a week**
    **Master''s and 10+ years of experience, Bachelor''s and 12+ years of experience or 18+ years in lieu of a degree**

    • Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction including:

      • Practical knowledge of Named Entity Recognition and Address Standardization to extract and clean unstructured text data.

      • Deep understanding of data matching strategies, including Blocking and Indexing, String Distance Metrics, and Phonetic Encoding.

      • Experience applying TD-IDF and Cosine Similarity for text comparisons and information retrieval.

    • Strong Python development skills for building analytics solutions and manipulating data.

    • Advanced SQL proficiency for complex data querying, optimization, and database operations.

    • Practical experience using Regex for advanced text processing, data cleansing, and pattern matching.

    • Familiarity with specialized libraries and frameworks including:

      • Linkage libraries such as Splink / FastLink, Dedupe, or recordlinkage

      • Core Python data science libraries, specifically spaCy for NLP tasks and Scikit-Learn for general machine learning and clustering

    • Familiarity with code reviews, version control, and maintaining data security and reproducibility standards.

    • Excellent communication skills.

      How To Stand Out From The Crowd (Desired Skills)

      • Prior experience delivering IT or data initiatives within federal, state, or local government environments

      • Proven ability to operate independently, take ownership of data pipeline architectures, and drive projects from data discovery through to post-implementation.

      • Experience retrieving, migrating, and manipulating data from legacy and distributed systems, including PostgreSQL, DB2, Oracle, SQL Server, and Hadoop, as well as unstructured flat files.

      • Experience utilizing Jenkins to automate continuous integration, testing, and deployment (CI/CD) for data validation pipelines.

      • Experience with pipeline automation tools to schedule and monitor complex data cleansing jobs.

      • Strong ability to translate complex algorithmic decisions (such as probabilistic match thresholds) into clear business logic for executive leadership and non-technical stakeholders.

      • Excellent problem-solving skills and proven verbal/written communication skills when collaborating across cross-functional teams.

Skills

Oracle
SQL
SQL Server
Machine Learning
NLP
Scikit-learn
Data Pipeline
Hadoop
Jenkins
PostgreSQL
Python

Similar jobs