Get more replies from employers
Send a job-specific resume in minutes.
TheCorporate LLC seeks an experienced Data Scientist with 10+ years of IT experience to lead entity resolution, record linkage, NLP, Python, and SQL initiatives. The role focuses on building scalable data processing, deduplication, and privacy-conscious data solutions.
This 100% onsite Woodlawn, MD position requires a Public Trust eligibility and hands-on work with big datasets, modern NLP techniques, and production-grade data pipelines.
Location: Woodlawn, MD
Work Model: 100% Onsite — 5 Days/Week
Employment Type: W2
Experience: 10+ Years
Compensation: $80–$100/hour W2
Clearance: Public Trust Required / Must Be Eligible to Obtain
We are seeking an experienced Data Scientist with 10+ years of IT experience and strong expertise in Entity Resolution, Record Linkage, NLP, Python, and Advanced SQL.
The ideal candidate will have hands-on experience building data processing and record-matching solutions that identify, link, cleanse, and deduplicate records across large and complex datasets.
Design, develop, and maintain advanced data processing and entity resolution pipelines using Python and SQL.
Apply NLP and information extraction techniques including Named Entity Recognition (NER), blocking and indexing, string distance metrics, TF-IDF, cosine similarity, phonetic encoding, and address standardization.
Clean, transform, validate, and manage large-scale datasets from diverse sources.
Develop solutions focused on data quality, integrity, security, and privacy.
Write and optimize complex SQL queries and database operations for performance and scalability.
Perform data validation, testing, deployment, and post-implementation monitoring.
Participate in code reviews and follow version-control and reproducibility standards.
Develop and maintain reliable, scalable data pipelines.
Translate complex matching algorithms, thresholds, and probabilistic scoring into understandable business logic.
Communicate technical findings and recommendations effectively to both technical and non-technical stakeholders.
10+ years of IT industry experience.
Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, Information Science, or a related field.
Strong hands-on experience with Entity Resolution, Record Linkage, and Deduplication.
Deep practical knowledge of NLP and information extraction.
Experience with:
Named Entity Recognition (NER)
Blocking and indexing
String distance metrics
TF-IDF
Cosine similarity
Phonetic encoding
Address standardization
Hands-on experience with Splink, FastLink, Dedupe, or recordlinkage.
Strong Python programming and data-pipeline development skills.
Advanced SQL skills, including complex queries and optimization.
Strong Regex skills for data cleansing and pattern matching.
Experience with spaCy and scikit-learn.
Strong written and verbal communication skills.
Experience working with federal, state, or local government data.
Experience with PostgreSQL, DB2, Oracle, SQL Server, Hadoop, and flat files.
Experience with Jenkins and CI/CD.
Experience with data pipeline scheduling, automation, and monitoring tools.
Experience owning data pipeline architecture from data discovery through production and post-implementation.
The selected candidate must be eligible to obtain and maintain a Public Trust suitability determination and successfully complete the required federal background investigation.
Previous Public Trust or federal clearance/suitability experience is highly preferred.
This is a 100% onsite position in Woodlawn, Maryland, Monday through Friday.
The position is not remote or hybrid. Candidates must currently live within reasonable commuting distance or have a firm relocation plan before starting the position.
Relocation assistance is not provided.
W2 employment only
Must be legally authorized to work in the United States
Must be able to work on W2 without sponsorship or transfer