Software Engineer, AI Evaluation

Nuna Inc.

San Francisco (CA)

On-site

USD 180,000 - 270,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nuna Inc. is building an AI health coach designed to be available 24/7, with evaluation at the core of safety and reliability.

You will own the end-to-end evaluation infrastructure, including harnesses, datasets, judges, and release gates, ensuring that the coach remains safe as it ships to patients. You'll work with data science, clinicians, and engineers to define standards, methodologies, and metrics, and you’ll implement tools for labeling, review workflows, and calibration to human labels

Qualifications

  • Extensive experience building and shipping reliable production systems.
  • Deep understanding of evaluating AI systems and agentic evaluation.
  • Ability to deploy evals and automated AI tooling in production.
  • Testing mindset for measurement systems, with coverage and regression discipline.
  • Ability to design workflows and UI for clinicians and labelers.

Responsibilities

  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling
  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners
  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients
  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop
  • Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding

Skills

Production systems
AI evaluation
Testing mindset
AI tooling in production
Statistics
UI for clinicians and labelers

Tools

LangSmith
Braintrust
DeepEval
Ragas
Promptfoo

Job description

Chronic disease isn\'t managed in a clinic. It is managed at home, in relationships, in the everyday. What\'s on the dinner table, what gets said, and who notices when someone\'s struggling. For the 130 million Americans managing a chronic condition, the healthcare system has offered the same answer for decades: a 15-minute doctor\'s visit, a pamphlet, and a portal login they\'ll never use.

At Nuna, we are building an AI health coach that shows up like a person who actually has time: available at 3am, infinitely patient, and never behind a waiting room. We use motivational interviewing to help patients and their families see themselves clearly, design experiments that fit their real lives, and navigate a system that has not historically been on their side. We are building from the ground up around a simple belief: patients don\'t want to be healthy; they want their lives back.

We\'re not competing with other health apps. We\'re competing with the moment a person gives up on getting better. If that\'s a problem you want to work on, we\'d like to talk.

Your team

We are a small, interdisciplinary team - engineers, data scientists, designers, product managers, and clinicians - building Nuna\'s AI health coach. Our products are only as good as the care and science behind them, and your piece is how we know the coach is safe and working. You\'ll own the evaluation system for the team building the coach: a data scientist partners with you on the science, clinicians and designers supply the ground truth, and the engineers shipping the agents depend on the signal you produce to decide what ships.

The role

This is a net-new, build-first role for someone who wants to own how we evaluate our AI agents end to end. You\'ll build the harnesses, datasets, judges, and release gates that tell us whether the coach is safe and good, and you\'ll own both that infrastructure and the evals that run on it. This is not a test-execution role - you write the code and own the system, rather than running tests someone else designed. You\'ll make the day-to-day calls on standards, methods, and trade-offs, often with incomplete information and the freedom to define the right answer yourself. Comfort with ambiguity is part of the job.

What you\'ll do

  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling
  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners
  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients
  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop
  • Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding
What we\'re looking for
  • Significant experience building and shipping reliable production systems and tooling
  • Deep understanding of how to evaluate AI systems - LLM-as-judge, red-teaming and adversarial testing, synthetic scenario generation, and multi-turn and agentic evaluation - and a clear sense of how evals themselves fail. You\'ve deployed evals and automated AI tooling in production, not just prototyped them
  • A testing mindset applied to building the measurement system, not running tests against a spec: adversarial instinct, coverage thinking, regression discipline, and documentation others can build on
  • You use AI in your daily work and build tools that make the people around you more effective
  • Enough fluency in statistics and experimental design to partner with a data scientist on calibration and reliability
  • Can design the workflow and build a functional UI for non-engineers like clinicians and labelers
  • A genuine interest in improving healthcare alongside an interdisciplinary team, with the judgment to tell a launch-blocking issue from a nice-to-have
Bonus Points
  • Experience in healthcare or another regulated, high-trust domain, and familiarity with the regulatory landscape
  • Hands-on experience with the eval tooling ecosystem (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar)
  • Red-teaming or AI safety experience - prompt injection, jailbreaks, adversarial and stress testing
  • Experience with automated, eval-driven model or prompt optimization
  • You\'ve built in an early-stage or fast-moving environment

Nuna is an Equal Employment Opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, disability, genetics and/or veteran status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
Software Engineer
Software Engineer

Nuna • San Francisco (CA)

Hybrid
USD 150,000 - 190,000
Senior Software Engineer
Senior Software Engineer

Apply • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer
Senior Software Engineer

Nuna • San Francisco (CA)

On-site
USD 120,000 - 150,000
Software Engineer
Software Engineer

Nuna Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
Lead Software Engineer
Lead Software Engineer

Nuna • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Evaluation & Safety Engineer
AI Evaluation & Safety Engineer

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
Founding AI Engineer, Agents & Evaluation
Founding AI Engineer, Agents & Evaluation

Zingage • New York (NY)

On-site
USD 190,000 - 240,000
Competitive base and meaningful equity
Equipment stipend
Luxury gym membership in NYC
+4
Senior AI Engineer – Agents
Senior AI Engineer – Agents

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
AI Safety & Evaluation Engineer — Health Tech
AI Safety & Evaluation Engineer — Health Tech

Nuna Inc. • San Francisco (CA)

On-site
USD 180,000 - 270,000