Applied AI Engineer — Evaluation & Benchmarking

Sable

San Francisco (CA)

On-site

USD 140,000 - 190,000

Full time

10 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Sable is seeking an engineer to translate performance signals into clear evaluation criteria for Aidan, the AI agent that leads customer calls. You will design experiments across voice, vision, and browser interactions, and drive improvements to the agent, context, and models in a real product environment.

You will build benchmarks and simulations, score customer-specific scenarios, and identify gaps in capabilities to accelerate the next generation of intelligent assistants.

Qualifications

  • Engineer with statistics and/or applied machine learning experience.
  • Comfortable in a production codebase.
  • Bonus: prior experience with evaluations or with voice, realtime, or browser use/computer use agents.

Responsibilities

  • Design evaluations for voice, vision, and browser use to direct and accelerate improvements in the agent, context, and models.
  • Build and maintain our internal benchmarks and simulations.
  • Work on systems to score customer-specific scenarios and identify gaps in current capabilities.

Skills

Machine learning
Statistics
Production codebase

Job description

Sable is seeking an engineer to translate performance signals into clear evaluation criteria for Aidan, the AI agent that leads customer calls. You will design experiments across voice, vision, and browser interactions, and drive improvements to the agent, context, and models in a real product environment.

You will build benchmarks and simulations, score customer-specific scenarios, and identify gaps in capabilities to accelerate the next generation of intelligent assistants.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied AI Engineer, Evals
Applied AI Engineer, Evals

Sable • San Francisco (CA)

On-site
USD 140,000 - 190,000
Applied AI Engineer, Generalist
Applied AI Engineer, Generalist

Sable • San Francisco (CA)

On-site
USD 180,000 - 240,000
Multimodal AI Agent Engineer - End-to-End Systems
Multimodal AI Agent Engineer - End-to-End Systems

Sable • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Deployment & Growth Engagement Lead
AI Deployment & Growth Engagement Lead

Sable AI Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
AI Evaluations Engineer — Benchmarking Frontiers
AI Evaluations Engineer — Benchmarking Frontiers

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
AI Agent Engineer – Open Source & Eval
AI Agent Engineer – Open Source & Eval

Sail • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior AI Benchmark SME for Quant, Science & Corporate
Senior AI Benchmark SME for Quant, Science & Corporate

Lilt • United States

Remote
USD 90,000 - 130,000
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation
Voice AI Evaluation Engineer
Voice AI Evaluation Engineer

Aircall • San Francisco (CA)

On-site
USD 181,000 - 250,000
Competitive salary package & benefits
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

EPAM Systems Inc • United States

Remote
USD 140,000 - 190,000