Remote AI Evaluation Engineer

Planet Pharma

San Francisco (CA)

Remote

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Planet Pharma is seeking experienced scientists and engineers to rigorously evaluate frontier AI models against real technical work, including data analysis, design verification, and experiment readouts.

You'll craft realistic tasks from your practice, run them through frontier AI agents, and judge outputs to professional standards, with emphasis on clear reasoning and reproducible results. This is a fully remote role across STEM and Data Science.

Qualifications

  • 2+ years of applied experience preferred in one of the core sciences or engineering disciplines or data science.
  • In progress Bachelor’s degree or higher; 2+ years experience outside undergrad is preferred.
  • Coding / scientific computing: comfortable writing and verifying work in code (Python or similar) and using command line.
  • Hands-on with real data and tools: instrument/test data, measurement files, simulations, schematics or drawings.
  • Familiarity with statistical reasoning: experiment design, measurement error and uncertainty.
  • Excellent written and spoken English; ability to articulate why a result is wrong.

Responsibilities

  • Design challenging, realistic technical tasks drawn from day-to-day work and author supporting files.
  • Run tasks through frontier AI models and evaluate outputs against professional standards.
  • Compare model outputs, determine better performance, and document gaps.
  • Create detailed grading rubrics and explain why a result passes or fails.
  • Flag concrete failures with evidence and ensure proper interpretation of the ask.
  • Review and refine tasks across disciplines and collaborate with other experts.

Skills

2+ years experience
Python
Data analysis
Experiment design
Statistical reasoning
English proficiency
Technical writing

Education

Bachelor's degree or higher

Tools

MATLAB
Python
R
SPICE
PCB/ circuit tools
FEA/CFD

Job description

Planet Pharma is seeking experienced scientists and engineers to rigorously evaluate frontier AI models against real technical work, including data analysis, design verification, and experiment readouts.

You'll craft realistic tasks from your practice, run them through frontier AI agents, and judge outputs to professional standards, with emphasis on clear reasoning and reproducible results. This is a fully remote role across STEM and Data Science.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Validation Scientist – Physics
Remote AI Validation Scientist – Physics

Planet Pharma • San Francisco (CA)

Remote
USD 60,000 - 110,000
AI Task Architect for Real-World STEM & Biology
AI Task Architect for Real-World STEM & Biology

Planet Pharma • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote AI Chemistry Task Designer and Evaluator
Remote AI Chemistry Task Designer and Evaluator

Planet Pharma • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote AI Legal Evaluator & Task Designer
Remote AI Legal Evaluator & Task Designer

Planet Pharma • San Francisco (CA)

Remote
USD 140,000 - 190,000
Remote AI Mech Eng Task Designer & Evaluator
Remote AI Mech Eng Task Designer & Evaluator

Planet Pharma • San Francisco (CA)

Remote
USD 90,000 - 150,000
Remote Clinical AI Evaluator & Task Designer
Remote Clinical AI Evaluator & Task Designer

Planet Pharma • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote AI Accounting Evaluator & Task Designer
Remote AI Accounting Evaluator & Task Designer

Planet Pharma • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote Senior AI Trainer & Evaluation Specialist
Remote Senior AI Trainer & Evaluation Specialist

Remotebridge • Northern (KY)

Hybrid
USD 67,000 - 112,000
Remote Electronics Engineer & AI Evaluation Expert
Remote Electronics Engineer & AI Evaluation Expert

DataAnnotation • Town of Texas (WI)

Remote
USD 55,000 - 172,000
Remote work
Flexible hours
Remote Bio AI Evaluation Scientist (PhD)
Remote Bio AI Evaluation Scientist (PhD)

Weekday 1 • United States

Remote
USD 83,000 - 124,000