Senior Machine Learning Engineer, Speech & LLM Training Data

Propio Language Services

Overland Park (KS)

On-site

USD 100,000 - 170,000

Full time

5 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Propio Language Services is seeking aSenior Machine Learning Engineer, Speech & LLM Training Data, to lead data pipelines for multilingual audio, annotation, and evaluation for our conversational AI systems. You will own end-to-end data workflows, from collection and curation to QA and dataset versioning, across our healthcare, legal, and enterprise customers.

The role blends hands-on ML engineering with data governance on AWS, using PyTorch, HuggingFace, and audio tools to build secure,

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, Electrical Engineering, Computational Linguistics, or related field, or equivalent practical experience.
  • 5+ years of experience in ML engineering, speech/audio ML, ML data engineering, NLP, or LLM training-data workflows.
  • Strong hands-on experience with Python, SQL, Linux, Git, and Docker.
  • Experience training or evaluating models using PyTorch, Hugging Face, or comparable ML frameworks.
  • Experience with FFmpeg and audio-processing libraries such as TorchCodec, torchaudio, librosa, or equivalent tools.
  • Experience with speech-processing tasks such as VAD, diarization, ASR, forced alignment, language identification, and audio-quality analysis.
  • Experience with Databricks/Spark, Parquet/Arrow, and large-scale dataset pipelines.
  • Working knowledge of AWS S3, SageMaker, Glue, Step Functions, IAM, and KMS.
  • Experience with an annotation platform such as Labelbox, Label Studio, Scale AI, Prodigy, Argilla, or custom internal tooling.
  • Experience with experiment tracking and data versioning tools such as MLflow, Weights & Biases, DVC, Delta Lake, or LakeFS.
  • Experience with multilingual speech, translation, annotation workflows, and evaluation datasets.

Responsibilities

  • Define the data roadmap for multilingual speech, translation, multimodal LLMs, and conversational AI.
  • Build audio-processing pipelines covering resampling, channel handling, VAD, diarization, language identification, transcription, alignment, and quality filtering.
  • Build dataset pipelines for cleaning, deduplication, PII/PHI redaction, quality scoring, sampling, balancing, versioning, and lineage.
  • Design annotation guidelines, QA rubrics, golden datasets, and reviewer workflows.
  • Build evaluation datasets, analyze model failures, and translate performance gaps into targeted data improvements.
  • Run training, fine-tuning, post-training, and evaluation experiments, including SFT, preference data, DPO/RLHF-style workflows, and synthetic data generation.
  • Productionize secure, traceable, and reproducible data and ML workflows on AWS.

Skills

Python
SQL
Linux
Git
Docker
PyTorch
Hugging Face
ML data engineering

Education

Bachelor’s or Master’s in CS/ML/Data Science/EE/Computational Linguistics

Tools

FFmpeg
torchaudio
librosa
Label Studio/Labelbox
MLflow/Weights & Biases

Job description

Description

Propio Language Services is one of the top 5 providers high-quality, real-time multilingual interpretation, translation, and localization services, operating at 9-figure scale across healthcare, legal, and other industries. We are driven by a passion for cutting-edge technology and exceptional service, building seamless experiences that bridge communication gaps across languages, cultures, and modalities.

Job Type

Full-time

Role

Propio is hiring a Senior Machine Learning Engineer, Speech & LLM Training Data to transform large volumes of multilingual conversational audio into high-quality training and evaluation datasets. This hands-on role owns audio processing, dataset curation, annotation and QA workflows, model training, and evaluation for our multilingual speech, translation, and conversational AI systems.

Key Responsibilities
  • Define the data roadmap for multilingual speech, translation, multimodal LLMs, and conversational AI.
  • Build audio-processing pipelines covering resampling, channel handling, VAD, diarization, language identification, transcription, alignment, and quality filtering.
  • Build dataset pipelines for cleaning, deduplication, PII/PHI redaction, quality scoring, sampling, balancing, versioning, and lineage.
  • Design annotation guidelines, QA rubrics, golden datasets, and reviewer workflows.
  • Build evaluation datasets, analyze model failures, and translate performance gaps into targeted data improvements.
  • Run training, fine-tuning, post-training, and evaluation experiments, including SFT, preference data, DPO/RLHF-style workflows, and synthetic data generation.
  • Productionize secure, traceable, and reproducible data and ML workflows on AWS.
Requirements
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, Electrical Engineering, Computational Linguistics, or a related field, or equivalent practical experience.
  • 5+ years of experience in ML engineering, speech/audio ML, ML data engineering, NLP, or LLM training-data workflows.
  • Strong hands-on experience with Python, SQL, Linux, Git, and Docker.
  • Experience training or evaluating models using PyTorch, Hugging Face, or comparable ML frameworks.
  • Experience with FFmpeg and audio-processing libraries such as TorchCodec, torchaudio, librosa, or equivalent tools.
  • Experience with speech-processing tasks such as VAD, diarization, ASR, forced alignment, language identification, and audio-quality analysis.
  • Experience with Databricks/Spark, Parquet/Arrow, and large-scale dataset pipelines.
  • Working knowledge of AWS S3, SageMaker, Glue, Step Functions, IAM, and KMS.
  • Experience with an annotation platform such as Labelbox, Label Studio, Scale AI, Prodigy, Argilla, or custom internal tooling.
  • Experience with experiment tracking and data versioning tools such as MLflow, Weights & Biases, DVC, Delta Lake, or LakeFS.
  • Experience with multilingual speech, translation, annotation workflows, and evaluation datasets.
Preferred Qualifications
  • Experience with multilingual telephony, healthcare, interpretation, or call-center audio.
  • Experience with tools such as Silero VAD, pyannote, WhisperX, NeMo, Kaldi, or equivalent speech technologies.
  • Experience with distributed processing or training using Ray, PySpark, or similar frameworks.
  • Experience with HIPAA, PHI/PII redaction, and secure data governance.
  • Experience with low-resource languages, accents, dialects, and code-switching.
  • Experience with synthetic data, active learning, weak supervision, or LLM-as-judge evaluation.
Notice of AI Use in Job Application Review

As part of our commitment in creating a fair, efficient, and consistent hiring process we may use artificial intelligence (AI) to help our recruiting teams organize, summarize, and analyze information provided by candidates, including resumes, application responses, and other materials submitted during the application process.AI may be used to identify patterns, highlight relevant skills, and experience, and assist in comparing a candidate’s qualifications with the requirement of a specific role. These tools are to improve efficiency and consistency while supporting more informed hiring decisions, which will ultimately be made by the hiring team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer, Speech & LLM Training Data
Senior Machine Learning Engineer, Speech & LLM Training Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 190,000
Senior ML Engineer, Multilingual Speech & LLM Data
Senior ML Engineer, Multilingual Speech & LLM Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 190,000
Applied Scientist/Research Engineer, LLM Training Data
Applied Scientist/Research Engineer, LLM Training Data

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 180,000
Senior ML Engineer, Speech & LLM Data Pipelines
Senior ML Engineer, Speech & LLM Data Pipelines

Propio Language Services • Overland Park (KS)

On-site
USD 100,000 - 170,000
Senior AI Engineer
Senior AI Engineer

Propio Language Services • Overland Park (KS)

On-site
USD 120,000 - 210,000
Senior Data Scientist
Senior Data Scientist

Ellipsis Health Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
401k matching
Flexible paid time off
Health insurance
Sr. Machine Learning Engineer, Speech LLM Evaluation
Sr. Machine Learning Engineer, Speech LLM Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer (LLM)
Machine Learning Engineer (LLM)

DeepRec.ai • Boston (MA)

Hybrid
USD 170,000 - 200,000
Senior AI Engineer
Senior AI Engineer

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 180,000
Lead Specialist, AI Scientist
Lead Specialist, AI Scientist

Pearson • Town of Poland (NY)

On-site
USD 90,000 - 130,000