Staff AI Evaluation Scientist

Visa Hunt

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 200,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Stock options
Health & wellness
Meals provided
Life & family
Vacation days
Sponsorship support
Team building

Job summary

Reflection Research Lab in San Francisco is seeking a candidate to conduct critical comparative analysis to advance our understanding of model capabilities. You will build and refine evaluation systems that create tight feedback loops between data, evals, and model behavior, and develop generalizable frameworks for reasoning, alignment, and usefulness.

Join a fast-moving startup to translate insights into model improvements, collaborate with pre-training, post-training, and applied teams, and

Qualifications

  • Strong statistical analysis and experimental design skills to rigorously measure model improvements.
  • Familiarity with LLM evaluation methodologies: static benchmarks, human preference evals, and/or agentic tasks.
  • High agency and thrive in a fast-paced startup environment; bias for impact over process.
  • Excited to work in a new frontier lab, defining how we measure and accelerate progress toward more capable models.
  • Collaborative, detail-oriented, and motivated by building the feedback loops that make models truly improve.

Responsibilities

  • Conduct critical comparative analysis to advance our understanding of model capabilities.
  • Build and refine evaluation systems and processes that create tight feedback loops between data, evals, and model behavior.
  • Develop generalizable evaluation frameworks that capture what matters for reasoning, alignment, and usefulness.
  • Collaborate closely with pre-training, post-training, and applied teams to translate insights into model improvements.
  • Push the boundaries of what’s measurable, from synthetic evals to human feedback and real-world interaction data.

Skills

Statistical analysis
LLM evaluation
Startup mindset
Evaluation frameworks
Collaboration

Job description

Reflection Research Lab in San Francisco is seeking a candidate to conduct critical comparative analysis to advance our understanding of model capabilities. You will build and refine evaluation systems that create tight feedback loops between data, evals, and model behavior, and develop generalizable frameworks for reasoning, alignment, and usefulness.

Join a fast-moving startup to translate insights into model improvements, collaborate with pre-training, post-training, and applied teams, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Evaluations Engineer — Open Foundation Models
Staff AI Evaluations Engineer — Open Foundation Models

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+3
Staff ML Evaluation & Data Engineer
Staff ML Evaluation & Data Engineer

Reflection • New York (NY)

On-site
USD 190,000 - 270,000
Stock options
Health & wellness benefits
Daily meals in office
+3
Member of Technical Staff - Evaluations
Member of Technical Staff - Evaluations

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+3
AI Evaluation Researcher: Post-Training Signals & Metrics
AI Evaluation Researcher: Post-Training Signals & Metrics

Mosaic.tech • San Francisco (CA)

On-site
USD 140,000 - 190,000
Member of Technical Staff - Evaluations
Member of Technical Staff - Evaluations

Reflection • New York (NY)

On-site
USD 100,000 - 150,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Chief AI Evaluation & Research
Chief AI Evaluation & Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
Applied AI Research Staff Engineer - On-Site SF
Applied AI Research Staff Engineer - On-Site SF

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 210,000
Member of Technical Staff - Evaluations
Member of Technical Staff - Evaluations

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Stock options
Health & wellness
Meals provided
+4
ML Research Engineer: Fine-Tuning & Evaluation Systems
ML Research Engineer: Fine-Tuning & Evaluation Systems

HonestAI • San Francisco (CA)

On-site
USD 140,000 - 190,000
Applied AI Researcher
Applied AI Researcher

Morpheus Talent Solutions • San Francisco (CA)

Hybrid
USD 200,000 - 350,000