Senior Applied Data Scientist | NDA

GT

Warszawa

On-site

PLN 180,000 - 260,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GT is seeking a Senior Applied Data Scientist to advance entity matching at scale. You will develop and test ML, embeddings, and LLM-based approaches for matching complex business records across multiple data sources.

The role emphasizes model quality, experimentation, and evaluation, with engineering partners helping productionize successful ideas. You'll explore foundation-model techniques to improve matching while keeping costs and latency in check, working closely with data and software

Qualifications

  • 5–8 years of relevant experience in Data Science or Applied Machine Learning.
  • Strong Python and SQL skills.
  • Experience with embeddings, semantic similarity, LLMs.
  • Hands-on training of supervised and unsupervised models.
  • Knowledge of neural networks and transformer architectures.
  • Proficiency with ML frameworks such as TensorFlow, PyTorch, PyCaret.
  • Experience retraining a taxonomy classifier in production.
  • Experimental judgment with baselines, metrics, and error analysis.
  • Ability to explain model behavior to engineering and business partners.

Responsibilities

  • Develop ML, embedding, and LLM-based approaches to matching entities.
  • Improve handling of messy data, including multilingual records.
  • Develop scoring and ranking to distinguish matches from duplicates.
  • Evaluate AI techniques balancing accuracy, scalability and cost.
  • Design scalable approaches considering model usage and compute cost.
  • Collaborate with data and software engineering to productionize models.
  • Define evaluation methods for match quality and benchmarks.
  • Build trusted benchmark sets for model comparison.
  • Explore LLM-assisted review for scalable benchmarking.

Skills

Applied Data Science
Python & SQL
Embeddings & LLMs
Neural networks & transformers
ML frameworks (TensorFlow, PyTorch, Py
Experimentation & evaluation
Engineer collaboration

Tools

TensorFlow
PyTorch
PyCaret

Job description

GT was founded in 2019 by a former Apple, Nest, and Google executive.GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.

On behalf of our client, GT is looking for a Senior Applied Data Scientist interested in developing and testing new ML, embedding, and LLM-based approaches to solve complex data matching problems at scale.

About the Client

Our client is a leading global management consultancy known for tackling some of the world’s most complex business challenges. With a focus on strategy, transformation, and performance improvement, the firm partners with major organizations across industries to drive lasting impact.

About the Role

We are looking for a Senior Applied Data Scientist to improve how entity resolution is performed at scale.

You will develop and test new ML, embedding, and LLM-based approaches for matching complex business records across multiple data sources.

The work is centered on model quality, experimentation, and evaluation; engineering partners will help productionize successful approaches.

A key part of the role is exploring how newer foundation-model techniques can improve matching quality while remaining practical and scalable for very large datasets.

Responsibilities:

Develop better ways to match company records

  • Build new ML, embedding, and LLM-based approaches for matching entities

  • Improve how the system handles messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies.

  • Develop scoring and ranking approaches to distinguish accurate matches from duplicates, similar-looking records, and unrelated entities.

  • Evaluate and implement AI and machine learning techniques to improve matching quality while considering accuracy, scalability, and cost.

  • Design approaches that can operate efficiently at scale, taking model usage and computational cost into consideration.

Improve evaluation, experimentation, and match quality

  • Define and improve methods for evaluating match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review effort.

  • Assist in building trusted benchmark sets that allow us to compare new models against the current matching engine before production rollout.

  • Explore LLM-assisted review and validation to assess matching performance and benchmark more scalable approaches.

  • Turn ambiguous matching problems into clear hypotheses, experiments, metrics, and recommendations.

Partner with engineering to bring successful ideas into production

  • Work closely with data engineering and software engineering teams to turn promising prototypes into production-ready matching logic.

  • Provide engineering partners with clear model specifications, evaluation results, expected behavior, edge cases, and rollout requirements.

  • Help determine the most appropriate matching techniques based on data characteristics, confidence levels, and cost considerations.

  • Continuously evaluate matching performance, investigate regressions, and recommend improvements to models and matching logic.

  • Clearly communicate technical tradeoffs related to matching performance, scalability, cost, latency, explainability, and operational considerations.

Essential knowledge, skills & experience:
  • 5–8 years of relevant experience in Data Science, Applied Data Science, Applied Machine Learning, or a similar role.

  • Strong applied ML fundamentals, with hands-on experience building and evaluating models on real data.

  • Excellent Python and SQL skills.

  • Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques.

  • Hands-on experience training supervised and unsupervised models, including classification and NLP tasks.

  • Working knowledge of neural network and transformer architectures.

  • Proficiency with common ML frameworks such as TensorFlow, PyTorch, and PyCaret.

  • Experience retraining a taxonomy classifier or maintaining classification models in production.

  • Experimental judgment: able to define baselines, metrics, test sets, and error analysis that show whether quality improved.

  • Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners.

Nice-to-have:
  • Experience with entity resolution, record linkage, deduplication, or similar matching problems.

  • Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation.

  • Experience applying LLMs or embeddings to business problems where cost and scale matter.

  • Exposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQuery.

  • Familiarity with company, domain, website, firmographic, or other business-entity data.

Interview Steps:
  1. GT interview with Recruiter

  2. Technical interview

  3. Final interview

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Applied Data Scientist - NDA
Senior Applied Data Scientist - NDA

gt-hq • Poland

On-site
PLN 180,000 - 280,000
Senior Applied Data Scientist: Entity Resolution at Scale
Senior Applied Data Scientist: Entity Resolution at Scale

GT • Warszawa

On-site
PLN 180,000 - 260,000
Senior Applied Data Scientist - Scalable Entity Matching
Senior Applied Data Scientist - Scalable Entity Matching

gt-hq • Poland

On-site
PLN 180,000 - 280,000
Senior/Lead Data Scientist (Toronto time zone)
Senior/Lead Data Scientist (Toronto time zone)

N-iX • Kraków

Hybrid
PLN 240,000 - 360,000
Flexible working format
Education reimbursement
Company events
Senior Data Scientist
Senior Data Scientist

SKMgroup • Województwo małopolskie

On-site
PLN 60,000 - 90,000
Large freedom and real influence
Team approach to challenges
Flexible working culture
+1
Senior Data Scientist
Senior Data Scientist

Tenth Revolution Group • Warszawa

On-site
PLN 260,000 - 380,000
Private medical care
Flexible benefits
Insurance
+3
Senior ML/AI Data Scientists IRC301293
Senior ML/AI Data Scientists IRC301293

GlobalLogic • Kraków

On-site
PLN 250,000 - 400,000
Senior Data Scientist
Senior Data Scientist

Billennium • Poland

On-site
PLN 180,000 - 240,000
Private healthcare
Multisport card
Life insurance
+2
Senior Data Scientist
Senior Data Scientist

Capgemini • Warszawa

Hybrid
PLN 180,000 - 260,000
Medicover health care
Multisport card
Capgemini Helpline therapy support
+1
Senior Data/ML Engineer Python, Spark, AWS, Glue, SageMaker, Lambda, TensorFlow, PyTorch, SQL W[...]
Senior Data/ML Engineer Python, Spark, AWS, Glue, SageMaker, Lambda, TensorFlow, PyTorch, SQL W[...]

Diverse CG Sp. z o.o. Sp.k. • Warszawa

Hybrid
PLN 180,000 - 240,000
Private medical care
Co-financing for the sports card
Dedicated consultant support