AI Safety & Evaluation Engineer — Health Tech

Nuna Inc.

San Francisco (CA)

On-site

USD 180,000 - 270,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nuna Inc. is building an AI health coach designed to be available 24/7, with evaluation at the core of safety and reliability.

You will own the end-to-end evaluation infrastructure, including harnesses, datasets, judges, and release gates, ensuring that the coach remains safe as it ships to patients. You'll work with data science, clinicians, and engineers to define standards, methodologies, and metrics, and you’ll implement tools for labeling, review workflows, and calibration to human labels

Qualifications

  • Extensive experience building and shipping reliable production systems.
  • Deep understanding of evaluating AI systems and agentic evaluation.
  • Ability to deploy evals and automated AI tooling in production.
  • Testing mindset for measurement systems, with coverage and regression discipline.
  • Ability to design workflows and UI for clinicians and labelers.

Responsibilities

  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling
  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners
  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients
  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop
  • Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding

Skills

Production systems
AI evaluation
Testing mindset
AI tooling in production
Statistics
UI for clinicians and labelers

Tools

LangSmith
Braintrust
DeepEval
Ragas
Promptfoo

Job description

Nuna Inc. is building an AI health coach designed to be available 24/7, with evaluation at the core of safety and reliability.

You will own the end-to-end evaluation infrastructure, including harnesses, datasets, judges, and release gates, ensuring that the coach remains safe as it ships to patients. You'll work with data science, clinicians, and engineers to define standards, methodologies, and metrics, and you’ll implement tools for labeling, review workflows, and calibration to human labels

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer for Health Coach Safety
AI Evaluation Engineer for Health Coach Safety

Nuna Inc • United States

Remote
USD 120,000 - 190,000
AI Evaluation & Safety Engineer
AI Evaluation & Safety Engineer

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
Lead Software Engineer — Build AI Health Coach Platform
Lead Software Engineer — Build AI Health Coach Platform

Nuna • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Engineer — AI Health Coach (SF Hybrid)
Software Engineer — AI Health Coach (SF Hybrid)

Nuna Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna Inc. • San Francisco (CA)

On-site
USD 180,000 - 270,000
Software Engineer - AI Health Coach, End-to-End Ownership
Software Engineer - AI Health Coach, End-to-End Ownership

Nuna • San Francisco (CA)

Hybrid
USD 150,000 - 190,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Software Engineer — AI Health Coach Platform
Senior Software Engineer — AI Health Coach Platform

Apply • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer — AI Health Coach Platform
Senior Software Engineer — AI Health Coach Platform

Nuna • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior AI Evaluation Engineer: AI Safety & Quality Lead
Senior AI Evaluation Engineer: AI Safety & Quality Lead

Description Ciklum • United States

On-site
USD 120,000 - 190,000