Quick Overview
Seniority
Mid Senior
Work mode
On Site
Location
Woodlawn, MD, United States
Posted
19 hours ago
OracleSQLSQL ServerMachine LearningNLPScikit-learnData PipelineHadoopJenkinsPostgreSQLPython
Job Description
PLEASE NOTE:
- IT IS 100 % On Site position in Woodlawn
- Selected candidate must be able to obtain and maintain a public trust clearance
- Selected candidate must be willing to work on-site in Woodlawn, MD 5 days a week
KEY REQUIRED SKILLS:
- Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction concepts including Named Entity Recognition, Blocking and Indexing, String Distance Metrics, TF-IDF/Cosine Similarity, Phonetics Encoding, Address Standardization
- Solid Python, Regex, and SQL experience
- Excellent Communication skills
POSITION DESCRIPTION:
- Develop Analytics Solutions: Design, implement, and maintain advanced data processing and entity resolution pipelines using Python and SQL on enterprise data platforms.
- Data Hygiene & Management: Clean, transform, and manage large-scale datasets from diverse, complex sources, ensuring absolute data integrity, reliability, and security.
- Performance Optimization: Optimize complex SQL queries and database operations to ensure efficient data access, processing, and scalability.
- Engineering Standards: Actively participate in code reviews, enforce version control, and uphold best practices for code quality, reproducibility, and data privacy.
- End-to-End Delivery: Support data validation, testing, deployment, and post-implementation monitoring in a fast-paced environment.
Requirements
BASIC QUALIFICATIONS:
- Master's and 10+ years of experience, Bachelor's and 12+ years of experience or 18+ years in lieu of a degree
- Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, or Information Science with experience in NLP, Text Processing, Information Extraction, Python, SQL, Regex and specialized libraries/frameworks.
- Overall 10+ years’ experience in IT industry
REQUIRED SKILLS:
- Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction including:
- Practical knowledge of Named Entity Recognition and Address Standardization to extract and clean unstructured text data.
- Deep understanding of data matching strategies, including Blocking and Indexing, String Distance Metrics, and Phonetic Encoding.
- Experience applying TD-IDF and Cosine Similarity for text comparisons and information retrieval.
- Strong Python development skills for building analytics solutions and manipulating data.
- Advanced SQL proficiency for complex data querying, optimization, and database operations.
- Practical experience using Regex for advanced text processing, data cleansing, and pattern matching.
- Familiarity with specialized libraries and frameworks including:
- Linkage libraries such as Splink / FastLink, Dedupe, or recordlinkage
- Core Python data science libraries, specifically spaCy for NLP tasks and Scikit-Learn for general machine learning and clustering
- Familiarity with code reviews, version control, and maintaining data security and reproducibility standards.
- Excellent communication skills.
DESIRED SKILLS:
- Prior experience delivering IT or data initiatives within federal, state, or local government environments
- Proven ability to operate independently, take ownership of data pipeline architectures, and drive projects from data discovery through to post-implementation.
- Experience retrieving, migrating, and manipulating data from legacy and distributed systems, including PostgreSQL, DB2, Oracle, SQL Server, and Hadoop, as well as unstructured flat files.
- Experience utilizing Jenkins to automate continuous integration, testing, and deployment (CI/CD) for data validation pipelines.
- Experience with pipeline automation tools to schedule and monitor complex data cleansing jobs.
- Strong ability to translate complex algorithmic decisions (such as probabilistic match thresholds) into clear business logic for executive leadership and non-technical stakeholders.
- Excellent problem-solving skills and proven verbal/written communication skills when collaborating across cross-functional teams.
Similar jobs
- LI
Senior Director, Analytics and Data Science, Core Product & Growth
NewLife360
Remote🇺🇸RemoteYesterdaySQLTableauCAD+4Technology - CH
Data Scientist (TS/SCI reqd) with Security Clearance
NewCommand Holdings a Pequot Company
Tampa, FL🇺🇸On-siteYesterdayDockerSQLMachine Learning+15Technology - MA
GCCS - Data Scientist (ACC/SG) with Security Clearance
Makai LLC
Hampton, VA🇺🇸$125k - $135k/yrOn-site6 weeks agoPL/SQLSQLT-SQL+5Technology - KE
Data scientist - Onsite in Dallas TX. Need only local profile. Must have strong experience with Python-based, deployed on KPaaS) leveraging Mixed Integer Programming (MIP).
NewKeylent
Dallas, TX🇺🇸Hybrid19 hours agoMongoDBSQLMachine Learning+1Technology - TE
Product Data Scientist
NewTechridge, Inc.
United States🇺🇸Hybrid19 hours agoSQLMachine LearningScrum+4Technology - SO
Sr. Data Scientist
NewSystem One
Bloomington, MN🇺🇸HybridYesterdayDockerFastAPIFlask+15Technology