Staff AI Evaluation Engineer for Open Models

Reflection AI Ltd

New York (NY)

On-site

USD 120,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided in office
Paid parental leave
Unlimited vacation (US) / 30 days (UK)
Visa sponsorship
Team off-sites and celebrations

Job summary

Reflection AI Ltd in New York is seeking a role focused on evaluating and advancing large language models. You will design experiments, build evaluation systems, and translate findings into model improvements, collaborating across pre-training and applied teams.

You have strong statistical design skills, familiarity with LLM evaluation methodologies, and thrive in a fast-paced startup. The role emphasizes impact, collaboration, and delivering measurable progress.

Qualifications

  • Strong statistical analysis and experimental design skills to rigorously measure model improvements.
  • Familiarity with LLM evaluation methodologies: static benchmarks, human preference evals, and/or agentic tasks.
  • High agency and thrive in a fast-paced startup environment; bias for impact over process.
  • Excited to work in a new frontier lab, defining how we measure and accelerate progress toward more capable models.
  • Collaborative, detail-oriented, and motivated by building the feedback loops that make models truly improve.

Responsibilities

  • Conduct critical comparative analysis to advance our understanding of model capabilities.
  • Build and refine evaluation systems and processes that create tight feedback loops between data, evals, and model behavior.
  • Develop generalizable evaluation frameworks that capture what matters for reasoning, alignment, and usefulness.
  • Collaborate closely with pre-training, post-training, and applied teams to translate insights into model improvements.
  • Push the boundaries of what’s measurable, from synthetic evals to human feedback and real-world interaction data.

Skills

Statistical analysis
Experimental design
LLM evaluation
High-ownership
Collaboration

Job description

Reflection AI Ltd in New York is seeking a role focused on evaluating and advancing large language models. You will design experiments, build evaluation systems, and translate findings into model improvements, collaborating across pre-training and applied teams.

You have strong statistical design skills, familiarity with LLM evaluation methodologies, and thrive in a fast-paced startup. The role emphasizes impact, collaboration, and delivering measurable progress.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Engineer — Open Models & RL Research
Staff ML Engineer — Open Models & RL Research

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Staff ML Engineer - Open-Models & RL Systems
Staff ML Engineer - Open-Models & RL Systems

reflectionai • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 300,000
Top-tier compensation
Stock options
Health & wellness
+5
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Customer-Facing ML Engineer: Fine-Tune Open Models
Customer-Facing ML Engineer: Fine-Tune Open Models

Reflection AI • New York (NY)

On-site
USD 140,000 - 210,000
Top-tier compensation
Stock options
Health & wellness: medical, dental, V
+5
Staff AI Evaluations Engineer — Open Foundation Models
Staff AI Evaluations Engineer — Open Foundation Models

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+3
AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home
Staff ML Evaluation & Data Engineer
Staff ML Evaluation & Data Engineer

Reflection • New York (NY)

On-site
USD 190,000 - 270,000
Stock options
Health & wellness benefits
Daily meals in office
+3
LLM Evaluation Scientist: Benchmarks & Failure Insights
LLM Evaluation Scientist: Benchmarks & Failure Insights

AI Chopping Block • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health coverage
Dental coverage
Vision coverage
+4
Senior AI/ML Engineer — LLM & RLHF Specialist
Senior AI/ML Engineer — LLM & RLHF Specialist

Prolific • Omaha (NE)

On-site
USD 100,000 - 140,000
Competitive pay rates
Flexible hours
Ability to work from home
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000