Get more replies from employers
Send a job-specific resume in minutes.
GT is seeking a Senior Applied Data Scientist to advance entity resolution at scale. You will develop and test ML, embedding, and LLM-based approaches for matching complex business records across data sources.
Collaboration with engineering will productionize successful methods, and you will explore foundation-model techniques for scalable improvement. You will evaluate match quality with precision/recall metrics, and drive experiments, baselines, and benchmarks while managing cost and latency
GT was founded in 2019 by a former Apple, Nest, and Google executive. GT's mission is to connect the world's best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.
On behalf of our client, GT is looking for a Senior Applied Data Scientist interested in developing and testing new ML, embedding, and LLM-based approaches to solve complex data matching problems at scale.
Our client is a leading global management consultancy known for tackling some of the world's most complex business challenges. With a focus on strategy, transformation, and performance improvement, the firm partners with major organizations across industries to drive lasting impact.
We are looking for a Senior Applied Data Scientist to improve how entity resolution is performed at scale.
You will develop and test new ML, embedding, and LLM-based approaches for matching complex business records across multiple data sources. The work is centered on model quality, experimentation, and evaluation; engineering partners will help productionize successful approaches. A key part of the role is exploring how newer foundation-model techniques can improve matching quality while remaining practical and scalable for very large datasets.
5-8 years of relevant experience in Data Science, Applied Data Science, Applied Machine Learning, or a similar role.
Strong applied ML fundamentals, with hands‑on experience building and evaluating models on real data.
Excellent Python and SQL skills.
Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques.
Hands‑on experience training supervised and unsupervised models, including classification and NLP tasks.
Working knowledge of neural network and transformer architectures.
Proficiency with common ML frameworks such as TensorFlow, PyTorch, and PyCaret.
Experience retraining a taxonomy classifier or maintaining classification models in production.
Experimental judgment: able to define baselines, metrics, test sets, and error analysis that show whether quality improved.
Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners.
Experience with entity resolution, record linkage, deduplication, or similar matching problems.
Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation.
Experience applying LLMs or embeddings to business problems where cost and scale matter.
Exposure to large‑scale data platforms such as Spark, Snowflake, Databricks, or BigQuery.
Familiarity with company, domain, website, firmographic, or other business‑entity data.