AI Engineer, Evals & Agent Quality

Town.com, Inc.

San Francisco (CA)

On-site

USD 140,000 - 220,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Location is San Francisco, CA; five days a week in our Financial District office. You’ll collaborate with engineers to instrument quality and drive the system from signal to fix.

Qualifications

  • Build and own eval systems to measure assistant quality across multi-step trajectories.
  • Develop golden datasets and labeling loops to keep benchmarks up to date.
  • Design and implement model routing to balance cost, quality, and speed.

Responsibilities

  • Build a generalized eval system that measures assistant quality across every surface it touches and across multi-step trajectories.
  • Stand up golden datasets and the labeling loop to keep benchmarks current and to validate improvements.
  • Build model routing and online evaluation tooling to steer to the best models.
  • Make every prompt and system change measurable so the team can move fast without breaking what works.
  • Partner with engineers across the product to instrument quality and close the loop from signal to fix.

Skills

LLM eval systems
Quality measurement
Model routing
Greenfield development

Job description

About Town

Town (town.com) is AI that starts from who you are. We build a persistent model of your identity, your voice, your judgment, your relationships, and your priorities, and use it to do real work on your behalf across every tool where you operate: email, calendar, documents, Slack, and more. Town doesn't wait for you to prompt it. It observes, learns, and acts. The more you use it, the more it becomes an extension of you.


Town was founded by Jean-Denis Greze (CEO), former CTO of Plaid, and Tony Vincent (CPO), former Director of Applied AI Product at Google. We're a small, talent-dense team backed by Andreessen Horowitz, Forerunner Ventures, First Round Capital, and Conviction, with more than $73M raised to date.


About the role

Town is building the most personalized, most capable AI assistant for everyone — one that knows you deeply, works across every tool you use, and gets sharper over time. Building the best assistant means proving it's the best: every model, prompt, and system change has to be measurably better, on every surface it touches.


That's what you'll own. You'll build the evals and quality systems that turn assistant performance into numbers the whole team can trust, measuring and improving the full multi-step trajectory the assistant takes to do real work. You'll build the model routing that puts the right model in the right place balancing cost, quality, and speed.


This is a foundational, 0→1 build with ownership to match: the eval framework, the golden datasets and labeling loop, model routing, and online measurement, and you set the bar for what "best" means at Town.


What you'll do


  • Build a generalized eval system that measures assistant quality across every surface it touches — and, crucially, across multi-step agent trajectories.


  • Stand up golden datasets and the labeling loop that keeps them up to date and constantly checking to validate improvements and avoid regressions.


  • Build model routing and online evaluation tooling to help us learn and route to the best models.


  • Make every prompt and system change measurable, so the team can move fast without breaking what works.


  • Partner with engineers across the product to instrument quality and close the loop from signal to fix.



You might thrive here if you...


  • Have built or owned LLM eval systems, or offline/online quality measurement at scale.


  • Think rigorously about measurement. Maybe that came from an MLE or applied-ML background, maybe not, the instinct for how to measure "better" matters more than the exact pedigree.


  • Know the eval landscape hands-on, off-the-shelf tooling and eval frameworks, and have opinions on what to reach for when.


  • Are comfortable reasoning about model routing and the tradeoffs between models.


  • Ship the fixes, not just the dashboards and metrics.


  • Are a senior or staff engineer comfortable in greenfield, where the system doesn't exist yet.



Location

San Francisco, CA. Five days a week in person at our Financial District office.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Product Engineer
AI Product Engineer

Town • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Backend Engineer
Staff Backend Engineer

Town • New York (NY)

On-site
USD 140,000 - 210,000
Product Support San Francisco, California · Full-time · $115K – $145K Apply
Product Support San Francisco, California · Full-time · $115K – $145K Apply

Town • San Francisco (CA), Northern (KY)

Hybrid
USD 90,000 - 130,000
Founding Customer Enablement
Founding Customer Enablement

Town • San Francisco (CA)

On-site
USD 90,000 - 130,000
AI Product Engineer San Francisco, California · Full-time · $225K – $300K
AI Product Engineer San Francisco, California · Full-time · $225K – $300K

Town • San Francisco (CA)

On-site
USD 120,000 - 160,000
Product Support
Product Support

Town • San Francisco (CA)

On-site
USD 70,000 - 110,000
In-office in SF
Staff Backend Engineer
Staff Backend Engineer

Town • San Francisco (CA)

On-site
USD 120,000 - 160,000
Founding Technical Recruiter
Founding Technical Recruiter

Town • San Francisco (CA)

On-site
USD 150,000 - 230,000
iOS Engineer
iOS Engineer

Town • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Backend Engineer San Francisco, California · Full-time · $250K – $300K
Staff Backend Engineer San Francisco, California · Full-time · $250K – $300K

Town • San Francisco (CA)

On-site
USD 120,000 - 180,000