Multilingual Data Engineer for Large-Scale LLM Pipelines

Reflection

New York (NY)

On-site

USD 140,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals in office
Parental leave (22 weeks)
Unlimited PTO (US)
Visa sponsorship
Team events

Job summary

Reflection in New York is seeking a researcher-engineer to design and operate multilingual data pipelines and advance our multilingual language model capabilities.

You will collaborate with pre-training, mid-training, and post-training teams to improve data quality, efficiency, and cultural fidelity across languages, while maintaining rigorous measurement and strong ownership in a fast-moving team. You will lead small projects and contribute to open foundational models.

Qualifications

  • Strong software engineering fundamentals and comfort processing web-scale datasets in distributed environments.
  • Experience building large-scale data pipelines for language models, machine translation, speech, or search — ideally covering more than one language.
  • Fluency or working proficiency in at least one language other than English, and genuine curiosity about how languages differ.
  • A rigorous, measurement-first mindset: you prove a data change helped rather than assuming it did.
  • Navigate trade-offs between research objectives and practical engineering realities.
  • Comfortable with ambiguity and high ownership in a small, fast-moving team. Passionate about advancing the frontier of intelligence.

Responsibilities

  • Design and operate large-scale multilingual data pipelines — sourcing, cleaning, deduplication, language identification, and script normalization across high- and low resource languages.
  • Define and enforce quality bars for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.
  • Design and run scientific experiments to advance our understanding of scaling large language models to improve multilingual data efficiency.
  • Lead small research projects independently while collaborating on larger initiatives.
  • Build evaluation sets and diagnostics that expose where model behavior degrades by language, register, or domain, and close those gaps with targeted data.
  • Work with pre-training, mid-training, and post-training teams to land measurable, step-function improvements in multilingual capability.

Skills

Software engineering
Data pipelines
Distributed systems
Multilingual NLP
Language proficiency
Ambiguity tolerance

Job description

Reflection in New York is seeking a researcher-engineer to design and operate multilingual data pipelines and advance our multilingual language model capabilities.

You will collaborate with pre-training, mid-training, and post-training teams to improve data quality, efficiency, and cultural fidelity across languages, while maintaining rigorous measurement and strong ownership in a fast-moving team. You will lead small projects and contribute to open foundational models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Multilingual Data Engineer & ML Researcher
Senior Multilingual Data Engineer & ML Researcher

Reflection AI • San Francisco (CA), New York (NY)

On-site
USD 150,000 - 210,000
Senior Multilingual Data Engineer - Scale Language Models
Senior Multilingual Data Engineer - Scale Language Models

Reflection • San Francisco (CA)

On-site
USD 140,000 - 210,000
Top-tier compensation and equity
Stock options for all contributors
Comprehensive health, dental, vision,
+5
Staff Data Engineer, Multilingual NLP & Open Models
Staff Data Engineer, Multilingual NLP & Open Models

Socket.dev • San Francisco (CA)

On-site
USD 130,000 - 170,000
Compensation
Stock options
Healthcare
+5
Multilingual LLM Data Engineer & Evaluation Scientist
Multilingual LLM Data Engineer & Evaluation Scientist

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 180,000
Staff Engineer, Multilingual Data & Language Pipelines
Staff Engineer, Multilingual Data & Language Pipelines

Visa Hunt • San Francisco (CA)

On-site
USD 120,000 - 180,000
Top-tier compensation
Stock options
Health & wellness
+3
Member of Technical Staff - Multilingual Data
Member of Technical Staff - Multilingual Data

Reflection • San Francisco (CA)

On-site
USD 140,000 - 210,000
Top-tier compensation and equity
Stock options for all contributors
Comprehensive health, dental, vision,
+5
Member of Technical Staff - Multilingual Data
Member of Technical Staff - Multilingual Data

Reflection • New York (NY)

On-site
USD 140,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+5
Applied Scientist/Research Engineer, LLM Training Data
Applied Scientist/Research Engineer, LLM Training Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 180,000
NLP ML Research Engineer - Multilingual Production Pipelines
NLP ML Research Engineer - Multilingual Production Pipelines

NLP PEOPLE • United States

Remote
USD 159,000 - 264,000
Employee stock options
Discretionary RSU awards
Discounted products
LLM Data Scientist: End-to-End Training & Evaluation
LLM Data Scientist: End-to-End Training & Evaluation

Propio Language Services • Overland Park (KS)

On-site
USD 120,000 - 180,000