Applied Scientist/Research Engineer, LLM Training Data

Propio

Overland Park (KS)

Hybrid

USD 120,000 - 180,000

Full time

38 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Propio Language Services is a top-five global language services provider delivering multilingual interpretation, translation, and localization at scale. We are hiring an Applied Scientist/Research Engineer, LLM Training Data to own data strategy, curation pipelines, annotation workflows, and evaluation datasets powering our multilingual AI systems.

This hands-on role requires building scalable data pipelines, designing QA processes, identifying model failure modes, and improving LLM performance

Qualifications

  • Master’s degree in CS, ML, or related field, or equivalent practical experience.
  • 4+ years of experience in AI data operations, NLP data engineering, or LLM data workflows.
  • Strong hands-on experience with Python, SQL, and dataset curation pipelines.
  • Experience with annotation workflows, QA rubrics, evaluation datasets, or human-in-the-loop data processes.
  • Familiarity with multilingual NLP, speech data, translation data, low-resource languages, or agentic AI datasets.
  • Working knowledge of AWS data/ML tools (SageMaker, Glue, S3, Bedrock, Lambda, IAM, KMS).

Responsibilities

  • Define the end-to-end data roadmap for multilingual and multimodal AI systems.
  • Design and build dataset curation pipelines for training, post-training, and evaluation.
  • Create annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows.
  • Build evaluation datasets and benchmarks, analyze model failure modes, and translate gaps into data improvements.
  • Support post-training data workflows such as SFT, instruction tuning, RLHF, and synthetic data generation.
  • Use modern annotation tools and AWS-based data infrastructure to scale secure AI data workflows.

Skills

Python
SQL
Data pipelines
Communication skills
NLP data

Education

Master's degree in CS/ML or related field
PhD in CS/ML or related field

Tools

SageMaker
Bedrock
S3
Glue
Lambda
IAM
KMS
EKS/ECS
Labelbox

Job description

Description

Propio Language Services is a top-five global language services provider and the fastest-growing company in the industry. Operating at nine-figure scale across healthcare, legal, and other sectors, Propio delivers high-quality, real-time multilingual interpretation, translation, and localization services. Driven by cutting-edge technology and exceptional service, we create seamless experiences that bridge communication gaps across languages, cultures, and communication channels.

Propio is hiring an Applied Scientist/Research Engineer, LLM Training Data to own the data strategy, curation pipelines, annotation workflows, and evaluation datasets that power our multilingual AI systems. This is a hands-on technical role for someone who understands how to manage the full AI data lifecycle, from acquisition, curation, annotation, and quality control to evaluation datasets and post-training data, to directly improve LLM performance. The ideal candidate can build scalable data pipelines, design high-quality annotation and QA processes, identify model failure modes, and close performance gaps through targeted data acquisition, curation, and synthetic data generation.

Key Responsibilities:
  • Define the end-to-end data roadmap for multilingual and multimodal AI systems, including text, speech, translation, interpretation, low-resource languages, and agentic AI workflows.
  • Design and build dataset curation pipelines for training, post-training, and evaluation, including cleaning, deduplication, filtering, PII redaction, quality scoring, sampling, balancing, and versioning.
  • Create annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows for multilingual, speech, translation, and vision data.
  • Build evaluation datasets and benchmarks, analyze model failure modes, and translate performance gaps into targeted data improvements.
  • Support post-training data workflows such as SFT, instruction tuning, preference data, RLHF/DPO-style data, reward model data, and synthetic data generation.
  • Use modern annotation tools and AWS-based data infrastructure to scale secure, traceable, and compliant AI data workflows.

Requirements

Qualifications:
  • Master’s degree in Computer Science, Machine Learning, Data Science, Computational Linguistics, Linguistics, Statistics, or a related field, or equivalent practical experience.
  • 4+ years of experience in AI data, ML data operations, NLP data engineering, applied ML, speech/translation data, or LLM data workflows.
  • Strong hands-on experience with Python, SQL, and dataset curation pipelines.
  • Experience with annotation workflows, QA rubrics, evaluation datasets, or human-in-the-loop data processes.
  • Familiarity with multilingual NLP, speech data, translation data, low-resource languages, conversational AI, or agentic AI datasets.
  • Working knowledge of AWS data and ML tools such as S3, Glue, SageMaker, Bedrock, Lambda, Step Functions, EKS/ECS, IAM, or KMS.
  • Strong communication skills and ability to work with ML engineers, applied scientists, product teams, linguists, data teams, and vendors.
Preferred Qualifications:
  • PhD in Computer Science, Machine Learning, NLP, Computational Linguistics, Data Science, Statistics, or a related field.
  • Experience with LLM post-training workflows such as SFT, instruction tuning, preference data, RLHF, DPO, reward modeling, or evaluation data generation.
  • Experience with synthetic data generation, active learning, weak supervision, LLM-as-judge workflows, or automated data quality scoring.
  • Experience with modern annotation and data platforms such as Labelbox, Scale AI, Prodigy, Argilla, Snorkel, Humanloop, or custom internal tooling.

#LI-JS1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer, Speech & LLM Training Data
Senior Machine Learning Engineer, Speech & LLM Training Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 190,000
Multilingual LLM Data Engineer & Evaluation Scientist
Multilingual LLM Data Engineer & Evaluation Scientist

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 180,000
Data Operations Lead
Data Operations Lead

Insight Global • United States

On-site
Medical insurance
Vision insurance
401(k)
Senior ML Engineer, Multilingual Speech & LLM Data
Senior ML Engineer, Multilingual Speech & LLM Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 190,000
NLP Data Scientist – LLM
NLP Data Scientist – LLM

Yakamconsulting • Northern (KY)

Hybrid
USD 120,000 - 190,000
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

Remote
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
Member of Technical Staff, Data
Member of Technical Staff, Data

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Data Engineer — AI/ML Data Platforms
Senior Data Engineer — AI/ML Data Platforms

Akoncagua AI • Lakeland (FL)

On-site
USD 120,000 - 180,000
Sr Data Scientist GenAI
Sr Data Scientist GenAI

Select Minds LLC • Dallas (TX)

Remote
USD 130,000 - 160,000
Competitive salary
Flexible schedule
Opportunity for advancement
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000