ML Engineer: Voice Models & Evaluation Architect

Aircall.io, Inc.

San Francisco (CA)

On-site

USD 181,000 - 250,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary package
Growth opportunities

Job summary

Aircall seeks a Machine Learning Engineer to build evaluation frameworks for AI agents across voice and chat products. You will design pipelines for evaluating TTS/ASR, analyze failure modes, and implement regression tests to ensure reliability as models evolve.

Collaboration with teams to establish metrics, datasets, and tooling for scalable, trustworthy evaluations. The role emphasizes hands-on model evaluation, data pipelines, and strong communication to explain complex concepts to

Qualifications

  • BS in Computer Science, Machine Learning, Statistics, or related field
  • 3+ years of experience in ML Engineering or Applied ML with 8+ years of overall experience
  • Strong experience in evaluating supervised, unsupervised, LLMs and deep learning models.
  • Hands-on experience in failure analysis and evaluating LLMs
  • Experience building automated evaluation systems
  • Strong communication skills to articulate complex technical concepts across technical and non-technical audiences
  • Hands-on experience training or fine-tuning voice/speech models (TTS, ASR, or speech-to-speech), including data pipeline construction and experimentation.

Responsibilities

  • Design and document comprehensive evaluation frameworks for Aircall’s AI agents across voice, chat and messaging.
  • Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data, iterating on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency.
  • Assess AI generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies.
  • Analyze system design decisions and identify strengths, weaknesses, and potential failure points.
  • Design annotation guidelines and workflows for human-labeled evaluation data, and calibrate LLM-as-judge systems against human raters to ensure automated evals stay trustworthy over time.
  • Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals in production, and flagging model or data drift before it impacts customers.
  • Own the metric contract for every published AI metrics, including definition, population, grain, rollup, validity window.
  • Build release gates, the offline regression suite each AI surface must pass before a prompt, model, or config change ships, measuring reliability across repeated trials, not just average pass rates.
  • Build voice-specific evaluation: simulated callers across accents, languages, background noise, barge-in, DTMF, and tool failures, with latency and ASR accuracy as first-class quality metrics.

Skills

MLOps / ML engineering
Model evaluation
Voice models (TTS/ASR)
Automated evaluation systems
Communication skills
Data pipelines
Failure analysis
LLM evaluation

Education

BS in Computer Science, Machine Learning, Statistics, or related field

Job description

Aircall seeks a Machine Learning Engineer to build evaluation frameworks for AI agents across voice and chat products. You will design pipelines for evaluating TTS/ASR, analyze failure modes, and implement regression tests to ensure reliability as models evolve.

Collaboration with teams to establish metrics, datasets, and tooling for scalable, trustworthy evaluations. The role emphasizes hands-on model evaluation, data pipelines, and strong communication to explain complex concepts to

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer: Voice Models & AI Evaluation Systems
ML Engineer: Voice Models & AI Evaluation Systems

Engg • San Francisco (CA)

On-site
USD 181,000 - 250,000
Voice AI Evaluation Engineer
Voice AI Evaluation Engineer

Aircall • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package & benefits
Machine Learning Engineer (Evals and Voice Models)
Machine Learning Engineer (Evals and Voice Models)

Aircall.io, Inc. • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package
Growth opportunities
Machine Learning Engineer (Evals and Voice Models)
Machine Learning Engineer (Evals and Voice Models)

Aircall • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package & benefits
Machine Learning Engineer (Evals and Voice Models)
Machine Learning Engineer (Evals and Voice Models)

Engg • San Francisco (CA)

On-site
USD 181,000 - 250,000
Staff AI Systems Engineer - LLM, Evaluation, Equity
Staff AI Systems Engineer - LLM, Evaluation, Equity

Maven • San Jose (CA)

On-site
USD 180,000 - 240,000
Physical Health Benefits
Mental Health Benefits
Emotional Health Benefits
+5
Sr. Machine Learning Engineer, Speech LLM Evaluation
Sr. Machine Learning Engineer, Speech LLM Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
ML Engineer
ML Engineer

Catalyst Labs • Illinois

On-site
USD 100,000 - 130,000
ML Engineer
ML Engineer

Catalyst Labs • Palo Alto (CA)

On-site
USD 100,000 - 140,000
ML Engineer
ML Engineer

Catalyst Labs • New Jersey

On-site
USD 100,000 - 150,000