Clinical AI Systems Expert - MD

Mercor

New York (NY)

On-site

USD 100,000 - 170,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor seeks residency-trained physicians to join a non-clinical cohort evaluating and developing clinical AI systems. You will apply clinical judgment to grading criteria, dialogue evaluation, and structured annotation to measure model performance.

The role is shared across streams, with the possibility of moving between projects as priorities shift. A strong clinical background and clear, structured reasoning are essential.

Qualifications

  • MD or DO with completed residency and active medical license.
  • 3+ years post-residency clinical experience.
  • Fluent written and spoken English.

Responsibilities

  • Grading-criteria development for models.
  • Clinical dialogue evaluation for accuracy and safety.
  • Clinical reasoning annotation describing differential reasoning.
  • Output review to flag hallucinations or unsafe advice.
  • Guideline authoring to standardize care.
  • Difficult-case writing to probe model limits.

Education

MD or DO with completed residency

Job description

About the Role

Mercor is hiring residency-trained physicians across specialties for non-clinical work developing and evaluating clinical AI systems. You will apply your clinical judgment to grading-criteria development, dialogue evaluation, and structured annotation work that determines how these systems are measured.

This is a non-clinical role — no direct patient care, and no responsibility for live diagnosis.

This is a shared expert pool. After onboarding you may be matched to any of several concurrent clinical workstreams based on your specialty, availability, and interest. You are not committing to a single project, and you may move between streams as priorities shift.

What you may work on

Work varies by workstream and may include:

  • Grading criteria development — taking a clinical question and breaking the ideal answer into discrete, checkable criteria, so a model response can be graded consistently rather than impressionistically.

  • Clinical dialogue evaluation — reviewing multi-turn clinical conversations and judging them for accuracy, safety, completeness, appropriate hedging, and whether escalation advice was correct.

  • Clinical reasoning annotation — recording how you would work through a case, including the differential you considered and rejected, not only the conclusion.

  • Output review — flagging hallucinated findings, dangerous omissions, unsupported certainty, and advice that is technically correct but clinically unsafe.

  • Guideline authoring — defining edge cases and standards of care for your specialty so annotation stays consistent across a large group of clinicians.

  • Difficult-case writing — constructing clinical questions that probe the limits of current model reasoning.

Task length varies by stream, from roughly 45 minutes for a dialogue evaluation up to an hour or more for grading-criteria authoring. You will get a specific throughput target for whichever stream you are matched to.

Required qualifications
  • MD or DO with a completed residency in any specialty

  • Active, unrestricted medical license in your country of practice

  • 3+ years post-residency clinical experience, practising or previously practising

  • Comfort writing structured clinical rationale that a non-specialist reviewer can follow

  • Written and spoken English fluency

  • Minimum 20 hours per week, with the ability to concentrate hours when a stream is time-boxed

Preferred qualifications
  • Board certification in your specialty

  • U.S. licensure, and familiarity with U.S. standards of care and clinical guidelines

  • Primary care, internal medicine, emergency medicine, or hospitalist background, where breadth of presentation matters most

  • Prior clinical annotation, AI evaluation, medical education, or question-writing experience

  • Grading-criteria design, resident assessment, or clinical guideline development experience

  • Published research or sustained technical writing (please link a sample)

Why this work

Most clinical AI failures are not exotic — they are ordinary questions answered with misplaced confidence. Catching that requires someone who has actually carried clinical responsibility. The standard these systems get held to is written by physicians, and here you would be writing it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Residency-Trained Physician - AI Trainer
Residency-Trained Physician - AI Trainer

Obsidian • Nashville (TN)

On-site
USD 130,000 - 170,000
General Clinician (MD/DO) - HLS Expert Pool
General Clinician (MD/DO) - HLS Expert Pool

Modern MedEd • Maryland

Hybrid
USD 120,000 - 180,000
Fully remote
Weekly payments via Stripe or Wise
Radiologist - AI Trainer
Radiologist - AI Trainer

Obsidian • Nashville (TN)

On-site
USD 120,000 - 180,000
Radiologist - AI Trainer
Radiologist - AI Trainer

Mercor • Nashville (TN)

On-site
USD 180,000 - 240,000
MD/DO Clinical AI Evaluator & Guideline Writer
MD/DO Clinical AI Evaluator & Guideline Writer

Dorado • United States

Remote
USD 83,000 - 165,000
Physician AI Trainer & Clinical Evaluation Specialist
Physician AI Trainer & Clinical Evaluation Specialist

Obsidian • Nashville (TN)

On-site
USD 130,000 - 170,000
Radiologist - Medical Imaging Specialist
Radiologist - Medical Imaging Specialist

Mercor • New York (NY)

On-site
USD 120,000 - 150,000
MD Physician — Clinical AI Evaluation & Reasoning
MD Physician — Clinical AI Evaluation & Reasoning

Mercor • New York (NY)

On-site
USD 100,000 - 170,000
Radiologist - Medical Imaging Specialist
Radiologist - Medical Imaging Specialist

Obsidian • New York (NY)

On-site
USD 170,000 - 250,000
Clinical Medicine Domain Expert
Clinical Medicine Domain Expert

Weekday AI (YC W21) • California City (CA)

Hybrid
USD 96,000 - 152,000