Socure is looking for a Staff Data Scientist to join the Watchlist Data Science team in San Francisco. This role focuses on entity matching and classification to support AML compliance, combining real-time NLP systems with technical leadership and measurable customer outcomes.
Role Overview
You will work on scaling watchlist matching capabilities, building NLP pipelines for named entity recognition and information extraction, and translating research into production systems. The work includes improving data inputs, consolidating identity representations, and supporting customer tuning of screening behavior through benchmarking and evaluation frameworks.
Key Responsibilities
- Improve the quality, coverage, and freshness of Watchlist data using next-generation ingestion pipelines.
- Design and run rigorous data quality analysis pipelines to identify anomalies, evaluate dataset health, and ensure high-fidelity inputs for downstream model training.
- Use NLP and AI to classify and enrich raw source data into normalized schemas, extracting structured entity attributes from sanctions, PEP, adverse media, and enforcement sources.
- Expand multilingual capabilities to support global screening across Latin and non-Latin scripts.
- Develop and enhance NLP systems for consolidating watchlist identity representations, including information extraction and NER pipelines for deduplication and alias resolution into canonical profiles.
- Develop approaches to model how entity profiles evolve over time as names, aliases, and sanctions status change.
- Measure and benchmark entity resolution quality to continuously improve coverage and accuracy.
- Design and scale advanced NLP models and algorithms for real-time name matching and identity classification across diverse, multilingual unstructured data sources.
- Build multi-signal risk scoring that unifies name similarity, entity type, geography, list type, and other attributes into calibrated risk scores.
- Maintain benchmarking frameworks, golden datasets, and regression tests to sustain high recall and precision for the match engine.
- Develop models and analytics to help customers tune screening thresholds to the right operating point based on risk appetite and entity mix.
- Create backtesting and counterfactual analysis capabilities to help customers and internal teams understand the impact of model and threshold changes on screening outcomes.
- Design evaluation frameworks for AI-powered autonomous decision systems, including correct behavior definition, confidence threshold calibration, and drift monitoring in production.
- Support Watchlist expansion into payment screening by performing mathematical analysis and feature engineering to detect AML risk patterns across transaction data and payment message fields.
- Develop and maintain the AML taxonomy and risk signal library that underpins Watchlist classification and detection.
- Apply graph-based methods to identify indirect risk exposure, including entities connected to sanctions risk even when not directly listed.
- Lead technical initiatives across Watchlist Data Science and shape the team’s long-term strategy for entity matching, enrichment, and AI.
- Collaborate with Product and Engineering to translate research into production-grade systems at scale.
- Stay current with advances in NLP, large language models, and entity resolution, prototyping and deploying relevant techniques for AML use cases (including advanced NER and LLM-based extraction).
- Mentor peers and contribute to a culture of technical rigor and continuous improvement.
Required Qualifications
- Master’s or PhD in Computer Science, Computational Linguistics, Statistics, Applied Mathematics, or a related field, or equivalent professional experience.
- 7+ years of experience in data science or machine learning with meaningful work in NLP, entity resolution, or information extraction.
- Strong preference for experience in AML, sanctions screening, adverse media, or financial crime detection.
- Hands-on experience building and deploying NLP pipelines for entity extraction, named entity recognition, and record linkage at production scale.
- Strong experience with multilingual NLP and non-Latin script processing (familiarity is a plus).
- Python proficiency and major ML library experience including PyTorch, spaCy, and HuggingFace Transformers.
- Strong SQL proficiency, along with experience with large-scale data pipelines and production ML systems.
- Excellent communication skills, including the ability to translate model performance tradeoffs into compliance and business language for non-technical audiences.
Preferred and Bonus Skills
- Experience with LLMs and agentic AI frameworks such as LangChain or LangGraph (plus).
Technologies
- Python, PyTorch, spaCy, HuggingFace Transformers
- SQL
- LangChain, LangGraph
- Natural Language Processing (NLP), Named Entity Recognition (NER), Information Extraction
- LLMs, large language models
- Graph-based methods
Location and Work Setup
- San Francisco, CA (onsite)
- Location requirement: you must be located in one of the talent hubs: New York, San Francisco, Seattle, or Miami.
Compensation
USD 191,000 - 230,000 per year.
Additional Notes
- Sponsorship: not available at this time.
Follow Socure: YouTube | LinkedIn | X (Twitter) | Facebook