AI Evaluation & Safety Engineer

Nuna

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nuna is building an AI health coach that helps patients and families regain their lives. We are hiring a role to own end‑to‑end evaluation of our AI agents, building harnesses, datasets, judges, and release gates that ensure safety and high quality before deployment.

You will collaborate with data scientists, clinicians, and engineers to design calibration, reliability metrics, and trusted evaluation pipelines.

Qualifications

  • Experience building and shipping production-grade evaluation systems.
  • Proven ability to evaluate AI systems with LLMs as judges and multi-turn prompts.
  • Experience deploying evals in production and building automated AI tooling.
  • Strong statistics and experimental design foundations.
  • Ability to design workflows and UI for clinicians and labelers.
  • Ability to work with ambiguous requirements and define release criteria.

Responsibilities

  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling.
  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners.
  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients.
  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop.
  • Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding

Skills

Production systems
Eval infrastructure
AI safety
Adversarial testing
AI eval tooling
UI for non-engineers

Tools

Python
LangSmith
Prompt engineering

Job description

Nuna is building an AI health coach that helps patients and families regain their lives. We are hiring a role to own end‑to‑end evaluation of our AI agents, building harnesses, datasets, judges, and release gates that ensure safety and high quality before deployment.

You will collaborate with data scientists, clinicians, and engineers to design calibration, reliability metrics, and trusted evaluation pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety & Evaluation Engineer — Health Tech
AI Safety & Evaluation Engineer — Health Tech

Nuna Inc. • San Francisco (CA)

On-site
USD 180,000 - 270,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna Inc. • San Francisco (CA)

On-site
USD 180,000 - 270,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
Software Engineer - AI Health Coach, End-to-End Ownership
Software Engineer - AI Health Coach, End-to-End Ownership

Nuna • San Francisco (CA)

Hybrid
USD 150,000 - 190,000
Software Engineer — AI Health Coach (SF Hybrid)
Software Engineer — AI Health Coach (SF Hybrid)

Nuna Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
Lead Software Engineer — Build AI Health Coach Platform
Lead Software Engineer — Build AI Health Coach Platform

Nuna • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Software Engineer — AI Health Coach Platform
Senior Software Engineer — AI Health Coach Platform

Apply • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Evaluation & Quality Engineer
AI Evaluation & Quality Engineer

Capitalrx • Charlotte (NC)

On-site
USD 134,800 - 168,500
Senior Software Engineer — AI Health Coach Platform
Senior Software Engineer — AI Health Coach Platform

Nuna • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Infrastructure Engineer, Closed-Loop Evaluation
AI Infrastructure Engineer, Closed-Loop Evaluation

Jobzhr • Mountain View (CA)

Hybrid
USD 194,000 - 352,000