Machine Learning Engineer — Multilingual Data

Featherless AI

Poland

On-site

PLN 60,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Meaningful equity
Real ownership over product
Work on global models

Job summary

An innovative AI company in Poland is searching for a Machine Learning Engineer to manage and enhance a multilingual data pipeline. This role involves collaborating with researchers to ensure robust model performance across languages. Applicants should have 3+ years of experience in ML, particularly with multilingual datasets, and a solid grasp of NLP principles. The company offers competitive compensation, emphasizing real ownership over impactful projects and opportunities to work globally beyond English-speaking regions.

Qualifications

  • 3+ years of experience as an ML Engineer, Applied Scientist, or similar role.
  • Strong experience working with multilingual datasets.
  • Solid understanding of NLP fundamentals.
  • Experience building scalable data pipelines.
  • Familiarity with Unicode and language-specific challenges.

Responsibilities

  • Design, build, and maintain multilingual datasets.
  • Develop data pipelines for collection, cleaning, and labeling.
  • Implement quality filters for dataset improvement.
  • Work with researchers on language coverage and benchmarks.
  • Analyze dataset bias and coverage gaps.
  • Support training workflows with quality data.

Skills

Experience with multilingual or non-English datasets
NLP fundamentals (tokenization, embeddings, language modeling)
Building scalable data pipelines (Python, Spark, Ray)

Job description

We’re looking for a Machine Learning Engineer to own and scale our multilingual data pipeline—from sourcing and curation to evaluation and continuous improvement. You’ll work closely with researchers and infra engineers to ensure our models perform robustly across languages, scripts, and cultural contexts.

This role sits at the intersection of data, research, and production ML and is ideal for someone who cares deeply about data quality, linguistic diversity, and model generalization beyond English.

What You’ll Do
  • Design, build, and maintain large-scale multilingual datasets across high- and low-resource languages

  • Develop data pipelines for collection, cleaning, normalization, deduplication, and labeling

  • Implement quality filters using statistical, heuristic, and model-based methods

  • Work with researchers to define language coverage, benchmarks, and evaluation metrics

  • Analyze dataset bias, coverage gaps, and failure modes across regions and scripts

  • Support training, fine-tuning, and distillation workflows with high-quality multilingual data

  • Continuously iterate on datasets based on model performance and real-world usage

What We’re Looking For
  • 3+ years of experience as an ML Engineer, Applied Scientist, or similar role

  • Strong experience working with multilingual or non-English datasets

  • Solid understanding of NLP fundamentals (tokenization, embeddings, language modeling)

  • Experience building scalable data pipelines (Python, Spark, Ray, or similar)

  • Familiarity with Unicode, scripts, tokenization challenges, and language-specific quirks

  • Comfort collaborating with researchers and translating research needs into production systems

Nice to Have
  • Experience with low-resource languages or multilingual benchmarks (e.g. FLORES, XTREME)

  • Exposure to LLM training, fine-tuning, or distillation

  • Linguistics background or experience working with native language experts

  • Contributions to open-source datasets or ML tooling

  • Experience with data quality evaluation at scale

Why Join
  • Real ownership over a core differentiator of the product

  • Work on models used globally, not just in English-speaking markets

  • Small, high-caliber team with deep ML and systems experience

  • Competitive compensation + meaningful equity at Series A stage

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Researcher – Multilingual Data
AI Researcher – Multilingual Data

Featherless AI • Poland

Remote
PLN 213,128 - 298,379
Competitive compensation
Meaningful equity
Access to large datasets and modern infrastructure
Multilingual Data ML Engineer – Pipelines & Quality
Multilingual Data ML Engineer – Pipelines & Quality

Featherless AI • Poland

Remote
PLN 60,000 - 90,000
Competitive compensation
Meaningful equity
Real ownership over product
+1
Machine Learning Engineer
Machine Learning Engineer

CMC Markets • Warszawa

On-site
PLN 180,000 - 250,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

NTIATIVE IT Recruitment • Warszawa

On-site
PLN 180,000 - 300,000
Senior Data Scientist
Senior Data Scientist

SKMgroup • Województwo małopolskie

On-site
PLN 60,000 - 90,000
Large freedom and real influence
Team approach to challenges
Flexible working culture
+1
Machine Learning Engineer
Machine Learning Engineer

NTIATIVE IT Recruitment • Warszawa

On-site
PLN 180,000 - 240,000
Applied Machine Learning Engineer
Applied Machine Learning Engineer

Hitachi Energy • Kraków

On-site
PLN 45,000 - 65,000
Competitive benefits for financial wellbeing
Support for physical and mental wellbeing
Senior ML Ops Engineer
Senior ML Ops Engineer

Internetwork Expert • Warszawa

On-site
USD 16,414 - 27,358
Fully remote position
Flexible working hours
Inspiring team of colleagues
+3
Lead ML Engineer
Lead ML Engineer

Internetwork Expert • Kraków

On-site
USD 16,414 - 32,829
Fully remote position
Flexible working hours
Inspiring global team
+3
Platform Engineer
Platform Engineer

Nearmap • Warszawa

On-site
PLN 180,000 - 320,000
Medical care
Sport Card
MultiLife