ML Engineer: Voice Models & AI Evaluation Systems

Engg

San Francisco (CA)

On-site

USD 181,000 - 250,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Aircall, an AI-powered customer communications platform, seeks an ML Engineer to lead evaluation of its AI agents across voice, chat, and messaging. You design evaluation frameworks, train voice models, and drive trustworthy AI metrics.

You will build automated pipelines, calibrate LLM-based judges, monitor production quality, and collaborate with cross-functional teams. A BS in CS/ML, 3+ years in ML engineering, 8+ years total, and hands-on experience with TTS/ASR are required.

Qualifications

  • BS in CS, ML, statistics or related field.
  • 3+ years in ML Engineering or Applied ML with 8+ years total.
  • Experience evaluating supervised, unsupervised, LLMs and deep learning models.
  • Hands-on failure analysis and evaluating LLMs.
  • Experience building automated evaluation systems.
  • Strong communication across technical and non-technical audiences.
  • Hands-on training or fine-tuning voice models (TTS, ASR) with data pipelines.

Responsibilities

  • Design and document evaluation frameworks for AI agents across voice, chat and messaging.
  • Train and fine-tune voice models using production and synthetic data.
  • Assess AI generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies.
  • Analyze system design decisions and identify strengths, weaknesses, and potential failure points.
  • Design annotation guidelines and workflows for human-labeled evaluation data, and calibrate LLM-as-judge systems.
  • Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals.
  • Own the metric contract for every published AI metrics, including definition, population, grain, rollup, validity window.
  • Build release gates, the offline regression suite each AI surface must pass before a prompt, model, or config change ships.
  • Build voice-specific evaluation: simulated callers across accents, languages, background noise, barge-in, DTMF, and tool failures, with latency and ASR accuracy as first-class quality metrics.

Skills

ML Engineering
Applied ML
Model evaluation
Failure analysis
Automated evaluation
Communication
Voice models
Data pipelines

Education

BS in Computer Science, ML, Statistics, or related field
MS/PhD in CS/ML/Statistics or related field

Job description

Aircall, an AI-powered customer communications platform, seeks an ML Engineer to lead evaluation of its AI agents across voice, chat, and messaging. You design evaluation frameworks, train voice models, and drive trustworthy AI metrics.

You will build automated pipelines, calibrate LLM-based judges, monitor production quality, and collaborate with cross-functional teams. A BS in CS/ML, 3+ years in ML engineering, 8+ years total, and hands-on experience with TTS/ASR are required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer: Voice Models & Evaluation Architect
ML Engineer: Voice Models & Evaluation Architect

Aircall.io, Inc. • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package
Growth opportunities
Voice AI Evaluation Engineer
Voice AI Evaluation Engineer

Aircall • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package & benefits
Machine Learning Engineer (Evals and Voice Models)
Machine Learning Engineer (Evals and Voice Models)

Aircall.io, Inc. • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package
Growth opportunities
Machine Learning Engineer (Evals and Voice Models)
Machine Learning Engineer (Evals and Voice Models)

Aircall • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package & benefits
Machine Learning Engineer (Evals and Voice Models)
Machine Learning Engineer (Evals and Voice Models)

Engg • San Francisco (CA)

On-site
USD 181,000 - 250,000
ML Engineer
ML Engineer

Catalyst Labs • New Jersey

On-site
USD 100,000 - 150,000
ML Engineer
ML Engineer

Catalyst Labs • Austin (TX)

On-site
USD 80,000 - 120,000
ML Engineer
ML Engineer

Catalyst Labs • New York (NY)

On-site
USD 90,000 - 130,000
ML Engineer
ML Engineer

Catalyst Labs • Sunnyvale (CA)

On-site
USD 90,000 - 140,000
ML Engineer
ML Engineer

Catalyst Labs • Illinois

On-site
USD 100,000 - 130,000