AI Engineer, Evals & Agent Quality

Town.com, Inc.

San Francisco (CA)

On-site

USD 250,000 - 300,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

You will run validation checks to confirm improvements, build model routing and online evaluation tooling, and ensure every prompt change is measurable. This role collaborates with product engineers to instrument quality and drive fixes.

Qualifications

  • Experience building or owning LLM eval systems.
  • Strong, rigorous approach to measurement design.
  • Hands-on familiarity with the eval landscape and tooling.
  • Comfort reasoning about model routing and tradeoffs.

Responsibilities

  • Design and implement a generalized eval system to measure assistant quality across all user touchpoints.
  • Ensure evaluation coverage includes multi-step agent trajectories, not just single-turn outputs.
  • Create and maintain golden datasets with a labeling loop to keep data current.
  • Run continuous validation checks to confirm improvements and avoid regressions.
  • Build model routing capabilities and online evaluation tooling to learn which models perform best.
  • Make every prompt and system change measurable to enable rapid iteration without breaking behavior.
  • Collaborate with product engineers to instrument quality and close the loop from signals to fixes.

Skills

LLM eval systems
Measurement design
Evaluation tooling
Dashboards & metrics

Tools

Eval frameworks

Job description

Town.com is building an AI assistant, and this role will own evaluation and agent quality end to end, from measurement to routing and regression prevention.

Responsibilities
  • Design and implement a generalized eval system to measure assistant quality across all user touchpoints
  • Ensure evaluation coverage includes multi-step agent trajectories, not just single-turn outputs
  • Create and maintain golden datasets along with a labeling loop to keep data current
  • Run continuous validation checks to confirm improvements and avoid regressions
  • Build model routing capabilities and supporting online evaluation tooling to learn which models perform best
  • Make every prompt and system change measurable so the team can iterate quickly without breaking working behavior
  • Collaborate with product engineers to instrument quality and close the loop from quality signals to implementation fixes
Requirements
  • Experience building or owning LLM eval systems, or delivering offline/online quality measurement at scale
  • Strong, rigorous approach to measurement design
  • Hands-on familiarity with the eval landscape, including off-the-shelf tooling and eval frameworks, plus the ability to choose approaches based on judgment
  • Comfort reasoning about model routing and tradeoffs between models
  • Track record of shipping fixes in addition to producing dashboards and metrics
  • Senior or staff-level capability operating in a greenfield environment where the evaluation system does not yet exist
Location
  • San Francisco, CA (onsite)
  • Five days a week in person at the Financial District office
Compensation
  • USD 250,000 - 300,000 per yearly
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Evaluation Architect & Model Routing Engineer
Senior AI Evaluation Architect & Model Routing Engineer

Town.com, Inc. • San Francisco (CA)

On-site
USD 250,000 - 300,000
AI Engineer, Evals & Agent Quality
AI Engineer, Evals & Agent Quality

Town • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Eval & Quality Engineer — Onsite SF
Senior AI Eval & Quality Engineer — Onsite SF

Town • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Senior AI Engineer - Agent Team
Senior AI Engineer - Agent Team

FurtherAI Inc • San Francisco (CA)

On-site
USD 120,000 - 150,000
Fully covered health, dental, and vision benefits
Competitive Compensation and stock options
Unlimited PTO
+4
Research Engineer - Evals
Research Engineer - Evals

Pantera Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash compensation
Equity opportunities
Relocation support
+1
AI Engineer – Onsite – San Francisco, CA
AI Engineer – Onsite – San Francisco, CA

RS Global Services • San Francisco (CA)

On-site
USD 150,000 - 230,000
Evals Lead
Evals Lead

Fluency Digital, Inc. • New York (NY)

On-site
USD 120,000 - 150,000
Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000
RL Environments Engineer
RL Environments Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000