Applied AI Engineer, Evals

Sable

San Francisco (CA)

On-site

USD 140,000 - 190,000

Full time

10 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Sable is seeking an engineer to translate performance signals into clear evaluation criteria for Aidan, the AI agent that leads customer calls. You will design experiments across voice, vision, and browser interactions, and drive improvements to the agent, context, and models in a real product environment.

You will build benchmarks and simulations, score customer-specific scenarios, and identify gaps in capabilities to accelerate the next generation of intelligent assistants.

Qualifications

  • Engineer with statistics and/or applied machine learning experience.
  • Comfortable in a production codebase.
  • Bonus: prior experience with evaluations or with voice, realtime, or browser use/computer use agents.

Responsibilities

  • Design evaluations for voice, vision, and browser use to direct and accelerate improvements in the agent, context, and models.
  • Build and maintain our internal benchmarks and simulations.
  • Work on systems to score customer-specific scenarios and identify gaps in current capabilities.

Skills

Machine learning
Statistics
Production codebase

Job description

About Sable

Sable built Aidan, the first AI employee who can lead customer calls using realtime voice, vision, and browser use. Aidan runs a live, two-way conversation inside a real product environment, clicking through the product like a human, watching the user's screen, and adapting the journey on the fly. Every conversation feeds a self-improving context graph we call the Brain, so Aidan gets smarter with each call.

The role

You translate what makes a good performance into clear, defensible signals, and use those signals to improve Aidan. Role involves carrying out evaluations of agents on a broad range of areas on the frontier including voice, vision, and browser use, and using the findings to optimize the agent, context, and models.

What You'll Do
  • Design evaluations for voice, vision, and browser use to direct and accelerate improvements in the agent, context, and models
  • Build and maintain our internal benchmarks and simulations
  • Work on systems to score customer-specific scenarios and identify gaps in current capabilities.
Who you are
  • Engineer with statistics and/or applied machine learning experience
  • Comfortable in a production codebase
  • Bonus: prior experience with evaluations or with voice, realtime, or browser use/computer use agents.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied AI Engineer, Generalist
Applied AI Engineer, Generalist

Sable • San Francisco (CA)

On-site
USD 180,000 - 240,000
Applied AI Engineer — Evaluation & Benchmarking
Applied AI Engineer — Evaluation & Benchmarking

Sable • San Francisco (CA)

On-site
USD 140,000 - 190,000
Engagement Lead
Engagement Lead

Sable AI Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Multimodal AI Agent Engineer - End-to-End Systems
Multimodal AI Agent Engineer - End-to-End Systems

Sable • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding Product Designer for AI Customer Experience
Founding Product Designer for AI Customer Experience

Sable • San Francisco (CA)

On-site
USD 100,000 - 160,000
Engagement Lead
Engagement Lead

Sable • New York (NY)

On-site
USD 120,000 - 190,000
Engagement Lead
Engagement Lead

Sable • San Francisco (CA)

On-site
USD 150,000 - 210,000
Founding Account Executive
Founding Account Executive

Sable • New York (NY)

On-site
USD 120,000 - 250,000
Software Engineer, Agent Platform
Software Engineer, Agent Platform

Sage Care • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Platform Engineer
Platform Engineer

Sable • San Francisco (CA)

On-site
USD 180,000 - 260,000