Senior Applied Data Scientist - NDA

gt-hq

Poland

On-site

PLN 180,000 - 280,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GT is seeking a Senior Applied Data Scientist to advance entity resolution at scale. You will develop and test ML, embedding, and LLM-based approaches for matching complex business records across data sources.

Collaboration with engineering will productionize successful methods, and you will explore foundation-model techniques for scalable improvement. You will evaluate match quality with precision/recall metrics, and drive experiments, baselines, and benchmarks while managing cost and latency

Qualifications

  • 5-8 years of relevant experience in Data Science or Applied ML.
  • Strong applied ML fundamentals with hands-on model evaluation on real data.
  • Excellent Python and SQL skills.
  • Experience with embeddings, semantic similarity, LLMs, or related AI techniques.
  • Hands-on experience training supervised and unsupervised models, including NLP tasks.
  • Knowledge of neural networks and transformer architectures.
  • Proficiency with ML frameworks such as TensorFlow, PyTorch, PyCaret.
  • Experience retraining taxonomy classifiers or production models.
  • Ability to define baselines, metrics, and error analyses to show quality improvements.
  • Ability to explain model behavior and tradeoffs to engineers and business partners.

Responsibilities

  • Develop and test ML, embedding, and LLM approaches for entity matching at scale.
  • Improve handling of messy data: multilingual records, aliases, domains, firmographics.
  • Develop scoring/ranking to distinguish accurate matches from duplicates.
  • Evaluate AI/ML techniques balancing accuracy, scalability, and cost.
  • Design scalable approaches considering model usage and compute costs.
  • Improve evaluation, experimentation, and match quality metrics.
  • Build benchmark sets to compare models before production rollout.
  • Explore LLM-assisted review to benchmark scalable solutions.
  • Turn ambiguous matches into clear hypotheses, experiments, metrics, and recommendations.
  • Collaborate with engineering to productionize successful ideas.
  • Work with data and software engineers to turn prototypes into production logic.
  • Provide model specs and evaluation results to stakeholders.
  • Determine appropriate matching techniques based on data characteristics and cost.
  • Continuously evaluate performance and investigate regressions.
  • Communicate tradeoffs related to performance, scalability, cost, and latency.

Skills

Python
SQL
Embeddings
LLMs
NLP
TensorFlow
PyTorch
PyCaret
Experimentation
Model deployment

Tools

TensorFlow
PyTorch
PyCaret
Spark
BigQuery

Job description

GT was founded in 2019 by a former Apple, Nest, and Google executive. GT's mission is to connect the world's best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.

On behalf of our client, GT is looking for a Senior Applied Data Scientist interested in developing and testing new ML, embedding, and LLM-based approaches to solve complex data matching problems at scale.

About the Client

Our client is a leading global management consultancy known for tackling some of the world's most complex business challenges. With a focus on strategy, transformation, and performance improvement, the firm partners with major organizations across industries to drive lasting impact.

About the Role

We are looking for a Senior Applied Data Scientist to improve how entity resolution is performed at scale.

You will develop and test new ML, embedding, and LLM-based approaches for matching complex business records across multiple data sources. The work is centered on model quality, experimentation, and evaluation; engineering partners will help productionize successful approaches. A key part of the role is exploring how newer foundation-model techniques can improve matching quality while remaining practical and scalable for very large datasets.

Responsibilities
  • Develop better ways to match company records
  • Build new ML, embedding, and LLM-based approaches for matching entities
  • Improve how the system handles messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies.
  • Develop scoring and ranking approaches to distinguish accurate matches from duplicates, similar-looking records, and unrelated entities.
  • Evaluate and implement AI and machine learning techniques to improve matching quality while considering accuracy, scalability, and cost.
  • Design approaches that can operate efficiently at scale, taking model usage and computational cost into consideration.
  • Improve evaluation, experimentation, and match quality
  • Define and improve methods for evaluating match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review effort.
  • Assist in building trusted benchmark sets that allow us to compare new models against the current matching engine before production rollout.
  • Explore LLM‑assisted review and validation to assess matching performance and benchmark more scalable approaches.
  • Turn ambiguous matching problems into clear hypotheses, experiments, metrics, and recommendations.
  • Partner with engineering to bring successful ideas into production
  • Work closely with data engineering and software engineering teams to turn promising prototypes into production‑ready matching logic.
  • Provide engineering partners with clear model specifications, evaluation results, expected behavior, edge cases, and rollout requirements.
  • Help determine the most appropriate matching techniques based on data characteristics, confidence levels, and cost considerations.
  • Continuously evaluate matching performance, investigate regressions, and recommend improvements to models and matching logic.
  • Clearly communicate technical tradeoffs related to matching performance, scalability, cost, latency, explainability, and operational considerations.
Essential knowledge, skills & experience

5-8 years of relevant experience in Data Science, Applied Data Science, Applied Machine Learning, or a similar role.

Strong applied ML fundamentals, with hands‑on experience building and evaluating models on real data.

Excellent Python and SQL skills.

Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques.

Hands‑on experience training supervised and unsupervised models, including classification and NLP tasks.

Working knowledge of neural network and transformer architectures.

Proficiency with common ML frameworks such as TensorFlow, PyTorch, and PyCaret.

Experience retraining a taxonomy classifier or maintaining classification models in production.

Experimental judgment: able to define baselines, metrics, test sets, and error analysis that show whether quality improved.

Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners.

Nice‑to‑have

Experience with entity resolution, record linkage, deduplication, or similar matching problems.

Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation.

Experience applying LLMs or embeddings to business problems where cost and scale matter.

Exposure to large‑scale data platforms such as Spark, Snowflake, Databricks, or BigQuery.

Familiarity with company, domain, website, firmographic, or other business‑entity data.

Interview Steps
  • GT interview with Recruiter
  • Technical interview
  • Final interview
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Applied Data Scientist | NDA
Senior Applied Data Scientist | NDA

GT • Warszawa

On-site
PLN 180,000 - 260,000
Senior Applied Data Scientist - Scalable Entity Matching
Senior Applied Data Scientist - Scalable Entity Matching

gt-hq • Poland

On-site
PLN 180,000 - 280,000
Senior Applied Data Scientist: Entity Resolution at Scale
Senior Applied Data Scientist: Entity Resolution at Scale

GT • Warszawa

On-site
PLN 180,000 - 260,000
Data Scientist
Data Scientist

Bain & Company • Warszawa

On-site
PLN 180,000 - 250,000
Senior Data Scientist
Senior Data Scientist

SKMgroup • Województwo małopolskie

On-site
PLN 60,000 - 90,000
Large freedom and real influence
Team approach to challenges
Flexible working culture
+1
Senior ML/AI Data Scientists IRC301293
Senior ML/AI Data Scientists IRC301293

GlobalLogic • Kraków

On-site
PLN 250,000 - 400,000
Senior Data/ML Engineer Python, Spark, AWS, Glue, SageMaker, Lambda, TensorFlow, PyTorch, SQL W[...]
Senior Data/ML Engineer Python, Spark, AWS, Glue, SageMaker, Lambda, TensorFlow, PyTorch, SQL W[...]

DCG Poland • Warszawa

Hybrid
PLN 180,000 - 300,000
Private medical care
Co-financing for sports card
Dedicated consultant support
Senior Data Scientist
Senior Data Scientist

Tenth Revolution Group • Warszawa

On-site
PLN 260,000 - 380,000
Private medical care
Flexible benefits
Insurance
+3
Senior Data/ML Engineer Python, Spark, AWS, Glue, SageMaker, Lambda, TensorFlow, PyTorch, SQL W[...]
Senior Data/ML Engineer Python, Spark, AWS, Glue, SageMaker, Lambda, TensorFlow, PyTorch, SQL W[...]

Diverse CG Sp. z o.o. Sp.k. • Warszawa

Hybrid
PLN 180,000 - 240,000
Private medical care
Co-financing for the sports card
Dedicated consultant support
Senior/Lead Data Scientist (Toronto time zone)
Senior/Lead Data Scientist (Toronto time zone)

N-iX • Kraków

Hybrid
PLN 240,000 - 360,000
Flexible working format
Education reimbursement
Company events