AI Evaluation Engineer for Health Coach Safety

Nuna Inc

United States

Remote

USD 120,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Nuna is building an AI health coach and seeks a data scientist to own the evaluation system end-to-end. You will build the harnesses, datasets, judges, and release gates that ensure safety and quality before any deployment, guiding what ships to patients.

You’ll collaborate with data scientists, clinicians, and engineers while defining standards, methods, and trade-offs in the face of incomplete information.

Responsibilities

  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling
  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners
  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients
  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios with

Job description

Nuna is building an AI health coach and seeks a data scientist to own the evaluation system end-to-end. You will build the harnesses, datasets, judges, and release gates that ensure safety and quality before any deployment, guiding what ships to patients.

You’ll collaborate with data scientists, clinicians, and engineers while defining standards, methods, and trade-offs in the face of incomplete information.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation & Safety Engineer
AI Evaluation & Safety Engineer

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
Software Engineer - AI Health Coach, End-to-End Ownership
Software Engineer - AI Health Coach, End-to-End Ownership

Nuna • San Francisco (CA)

Hybrid
USD 150,000 - 190,000
Lead Software Engineer - End-to-End AI Health Coach
Lead Software Engineer - End-to-End AI Health Coach

Nuna • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer — AI Health Coach (SF Hybrid)
Software Engineer — AI Health Coach (SF Hybrid)

Nuna Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Software Engineer — AI Health Coach Platform
Senior Software Engineer — AI Health Coach Platform

Apply • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer — AI Health Coach Platform
Senior Software Engineer — AI Health Coach Platform

Nuna • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Evaluation Scientist: Trust & Safety Leader
AI Evaluation Scientist: Trust & Safety Leader

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
AI Safety Evaluator for Real-World AI Systems
AI Safety Evaluator for Real-World AI Systems

mpathic • Seattle (WA)

On-site
USD 70,000 - 110,000
AI Safety & Behavioral Health Evaluator
AI Safety & Behavioral Health Evaluator

HumanitApp • Northern (KY)

Hybrid
USD 62,000 - 96,000