AI Researcher – Multilingual Data

Featherless AI

Poland

On-site

PLN 213,128 - 298,379

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Meaningful equity
Access to large datasets and modern infrastructure

Job summary

A pioneering AI startup in Poland is seeking an AI Researcher focused on multilingual data. The role involves designing research on multilingual datasets, developing strategies for low-resource languages, and publishing findings at respected conferences. Ideal candidates will have a strong background in NLP/ML research and proficiency in Python with frameworks like PyTorch. The position offers competitive compensation and real ownership over research direction in a fast-paced environment.

Qualifications

  • Strong background in multilingual NLP or ML research.
  • Experience with large-scale text datasets across multiple languages.
  • Comfortable prototyping in Python with modern ML frameworks.

Responsibilities

  • Design and execute research on multilingual datasets.
  • Develop strategies for low-resource and long-tail languages.
  • Publish research at top venues and contribute to open-source.

Skills

NLP / ML research
Prototyping in Python
Data quality metrics
Transfer learning

Education

Publication record at top conferences

Tools

PyTorch
JAX

Job description

About the Role

We’re looking for an AI Researcher focused on multilingual data to help us build and scale next-generation language models across diverse languages and domains. You’ll own research and execution around data sourcing, curation, evaluation, and training strategies for multilingual and low-resource languages, with a strong emphasis on publishing high-quality research and translating it into production systems.

This role is ideal for someone who enjoys working close to the frontier: balancing papers, prototypes, and real-world impact in a fast-moving startup environment.

What You’ll Do
  • Design and execute research on multilingual datasets, including data collection, filtering, deduplication, and quality measurement
  • Develop strategies for low-resource and long-tail languages (sampling, augmentation, curriculum design)
  • Research and improve cross-lingual transfer, alignment, and robustness in large language models
  • Build and maintain evaluation benchmarks for multilingual performance
  • Collaborate with engineers and researchers on training pipelines and model architecture decisions
  • Publish research at top venues (e.g., ACL, EMNLP, NeurIPS, ICML, ICLR) and contribute to open-source when appropriate
  • Translate research insights into practical improvements in production models
What We’re Looking For
  • Strong background in NLP / ML research, with a focus on multilingual or cross-lingual modeling
  • Publication record at respected conferences or journals (ACL, EMNLP, NeurIPS, ICML, ICLR, etc.)
  • Experience working with large-scale text datasets across multiple languages
  • Solid understanding of:
    • Tokenization and vocabulary design for multilingual models
    • Data quality metrics, filtering, and dataset bias
    • Transfer learning and multilingual representation learning
  • Comfortable prototyping in Python with modern ML frameworks (PyTorch, JAX, etc.)
  • Ability to operate independently and ship research in a startup pace environment
Nice to Have
  • Experience with low-resource languages or non-Latin scripts
  • Open-source contributions in NLP or data tooling
  • Experience training or evaluating large language models
  • Familiarity with multilingual benchmarks (e.g., XTREME, FLORES, TyDi QA)
Why Join Us
  • Real ownership over research direction and impact
  • A team that values papers and production
  • Access to meaningful scale: large datasets, modern infrastructure, and fast iteration
  • Competitive compensation and meaningful equity at an early stage
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer — Multilingual Data
Machine Learning Engineer — Multilingual Data

Featherless AI • Poland

Remote
PLN 60,000 - 90,000
Competitive compensation
Meaningful equity
Real ownership over product
+1
AI Researcher: Multilingual Data to Production
AI Researcher: Multilingual Data to Production

Featherless AI • Poland

Remote
PLN 213,128 - 298,379
Competitive compensation
Meaningful equity
Access to large datasets and modern infrastructure
AI Researcher — Training Optimization
AI Researcher — Training Optimization

Featherless AI • Poland

Remote
PLN 80,000 - 110,000
AI Researcher — AI Architecture Research
AI Researcher — AI Architecture Research

Featherless AI • Poland

Remote
PLN 298,379 - 383,631
Competitive compensation
Meaningful equity
High ownership over research direction
Senior Data Scientist
Senior Data Scientist

SKMgroup • Województwo małopolskie

On-site
PLN 60,000 - 90,000
Large freedom and real influence
Team approach to challenges
Flexible working culture
+1
AI Researcher — Distillation
AI Researcher — Distillation

Featherless AI • Poland

Remote
PLN 298,379 - 383,631
Real ownership over research direction
Strong support for publishing and open research
Access to meaningful compute and production-scale problems
ML Researcher (NXJ-72)
ML Researcher (NXJ-72)

Newxel • Poland

On-site
PLN 60,000 - 90,000
Competitive salary and benefits package
Medical insurance
Top equipment kit
+1
Staff Machine Learning Researcher
Staff Machine Learning Researcher

RTB House • Warszawa

Hybrid
PLN 350,000 - 550,000
Attractive compensation
Remote and on-site work
Flexible hours
+1
AI Scientist
AI Scientist

Mistral Ai • Warszawa

On-site
PLN 180,000 - 320,000
Lead/Senior Data Scientist (Generative AI)
Lead/Senior Data Scientist (Generative AI)

Lingaro • Poland

On-site
PLN 279,000 - 334,800
Stable employment
Office as an option
Workation options
+3