Founding Research Engineer

Nolla Health

New York (NY)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Meal stipends
Team retreats
Unlimited vacation
Health insurance

Job summary

Nolla Health in New York City is seeking a hands-on engineer to lead evaluation infrastructure for our AI medical models. You’ll design studies, run experiments, and publish benchmarks that help shape product direction and patient care.

You will work directly with the founding team and clinicians, turning real cases into evaluative data with provenance and bias controls, and you’ll help scale post-training rewards and data curation efforts as we grow.

Qualifications

  • Built eval infrastructure for LLM systems
  • Opinions about contamination, difficulty calibration, judge bias, and reward hacking
  • You can write a paper and publish empirical ML work
  • Excited to work with physicians and turn their judgment into model performance
  • Worked on small teams and built things from zero

Responsibilities

  • Design and maintain the benchmarks used to judge clinical accuracy, safety, and documentation quality
  • Extend benchmarks to the agentic system: tool use, long-horizon tasks, memory of patient history, labs, escalation behavior
  • Build the grading stack: deterministic checks, rubric graders, and LLM judges calibrated against clinician ratings
  • Improve the production harness the product runs on
  • Turn real, consented cases into eval data with clinician review, de-identification and provenance
  • Run clinician review workflows that produce rubrics, labels, and feedback at scale
  • Lead clinical evaluations from protocol through publication and publish benchmarks for the field
  • Collaborate with post-training on reward design and data curation, and run experiments where it helps

Skills

Eval infrastructure
Contamination risk awareness
Judgment bias awareness
Paper writing
Collaboration with physicians

Tools

Inspect framework
Verifiers
Harbor

Job description

The mission

Nolla is building AI doctors so that top-quality healthcare is accessible to everyone. Nolla Derm is the #1 medical skincare treatment app in the U.S. App Store. We've treated thousands of acne patients in the U.S., scanned 1% of Norway's population for skin cancer, and our in-house clinical models are state of the art on clinical benchmarks. We just launched NollaMD, our urgent care app, and we're building dedicated specialty apps for conditions like women’s health. We’ve raised $6.5M from General Catalyst and other strategic investors.

The role

We have built our own clinical benchmarks, post-trained on them, and produced models that beat frontier performance on clinical tasks. We are in a unique position of both delivering care to real patients and training the models that deliver it. You’ll work on our evals and harnesses, build the data loop with clinicians that feeds them, and lead where they go next.

You'll work directly with the founding team and with our post-training partners, and your work reaches patients immediately. This is a hands‑on role: you will write code most days, design studies, and be the author of record on what we publish.

What you'll build
Evals and benchmarks
  • Design and maintain the benchmarks we use to judge clinical accuracy, safety, and documentation quality

  • Extend them to our agentic system: tool use, long-horizon tasks that span many visits, memory of a patient's history, uploaded records, labs, and escalation behavior

Graders and harnesses
  • Build the grading stack: deterministic checks, rubric graders, and LLM judges calibrated against clinician ratings

  • Improve the production harness the product runs on

Data loop with clinicians
  • Turn real, consented cases into eval cases and training examples, with clinician review, de-identification, and provenance built in

  • Run the clinician review workflows that produce rubrics, labels, and feedback at scale

Research
  • Lead our clinical evaluations from protocol through publication, and publish our benchmarks for the field

  • Collaborate with post-training on reward design and data curation, and run experiments where it helps

What you bring
  • You have built eval infrastructure for LLM systems that other people depended on

  • You have opinions about contamination, difficulty calibration, judge bias, and reward hacking

  • You can write a paper. Authorship on empirical ML work, ideally with a human comparison or a benchmark release

  • You are excited to work with some of the best physicians in the country and turn their judgment into model performance

  • You've worked on small teams and built things from zero

Bonus points
  • Clinical AI evaluation experience: rubric-based health benchmarks, simulated-patient studies, agentic clinical benchmarks

  • Post-training experience (RLHF, DPO, GRPO or similar)

  • Experience with multimodal models

  • Familiarity with eval and environment frameworks such as Inspect, Verifiers, or Harbor

Role logistics, compensation & benefits
  • Role Type: Engineering

  • Salary: $180,000–$240,000, based on experience

  • Job Type: Full-time

  • Work Setup: In-person, New York City

  • Equity: Meaningful equity, commensurate with experience

  • Health Insurance: Medical, dental, and vision

  • HSA/FSA: Eligible

  • Time Off: Flexible, unlimited vacation

  • Additional Perks: Meal stipends, team retreats

  • Work Hours: Flexible but demanding. We're building something that matters

  • Growth: Founding team members step into expanded roles as we scale

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding AI Engineer
Founding AI Engineer

Nolla Health • New York (NY)

On-site
USD 160,000 - 220,000
Health Insurance
Flexible unlimited vacation
Meal stipends
+1
Director of Machine Learning (Healthcare AI)
Director of Machine Learning (Healthcare AI)

Nxt Level • United States

On-site
USD 150,000 - 200,000
Competitive salary
Meaningful equity
Direct line to CEO
+1
Applied AI Engineer
Applied AI Engineer

Sarah Smith Fund • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Remote-first culture
Direct access to founders
Professional development budget
Full Stack Engineer
Full Stack Engineer

HealthLeap • San Francisco (CA)

On-site
USD 175,000 - 275,000
Unlimited PTO
401(k) with 4% match
Laptop and home office budget
+1
Applied AI Engineer
Applied AI Engineer

Soulside, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Health insurance
Dental insurance
Vision insurance
+5
Senior AI Engineer
Senior AI Engineer

Doctronic • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Health benefits
Parental Leave: 12 weeks
Full Stack Engineer
Full Stack Engineer

HealthLeap AI • San Francisco (CA)

On-site
USD 175,000 - 275,000
100% healthcare premiums covered
Unlimited PTO + minimum recommended leave
401(k) with 4% match
+1
Applied AI Engineer
Applied AI Engineer

Norbert Health • New York (NY)

On-site
USD 100,000 - 150,000
Equity participation
Competitive salary
High autonomy and technical ownership
Backend Engineer
Backend Engineer

SupportFinity™ • San Francisco (CA)

On-site
USD 200,000 - 275,000
Equity ownership
Healthcare premiums covered
Unlimited PTO
+1
Full-Stack Engineer
Full-Stack Engineer

Precision Labs • Northern (KY), San Diego (CA)

Hybrid
USD 110,000 - 140,000
Stock options
Unlimited PTO
Medical, dental, and vision coverage
+1