Staff Evaluations Scientist for Open-Weight Models

Reflection AI

Greater London

On-site

GBP 75,000 - 110,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Comprehensive health insurance
Fully paid parental leave
Paid time off and relocation support
Daily lunch and dinner provided

Job summary

An AI research company in Greater London is seeking a candidate for a role focused on evaluating AI models. The position involves conducting comparative analysis, developing evaluation frameworks, and collaborating with various teams to improve model features. Ideal candidates will have strong statistical analysis skills and familiarity with LLM evaluation methodologies. The role is designed for those excited to thrive in a fast-paced startup environment and contribute to groundbreaking AI developments.

Qualifications

  • Strong statistical analysis and experimental design skills to rigorously measure model improvements.
  • Familiarity with LLM evaluation methodologies such as static benchmarks and human preference evals.
  • High agency and thrive in a fast-paced startup environment; bias for impact over process.

Responsibilities

  • Conduct critical comparative analysis to advance understanding of model capabilities.
  • Build and refine evaluation systems and processes for tight feedback loops.
  • Develop generalizable evaluation frameworks for reasoning, alignment, and usefulness.

Skills

Statistical analysis
Experimental design
LLM evaluation methodologies
Collaboration

Job description

An AI research company in Greater London is seeking a candidate for a role focused on evaluating AI models. The position involves conducting comparative analysis, developing evaluation frameworks, and collaborating with various teams to improve model features. Ideal candidates will have strong statistical analysis skills and familiarity with LLM evaluation methodologies. The role is designed for those excited to thrive in a fast-paced startup environment and contribute to groundbreaking AI developments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier AI Evaluation Scientist
Frontier AI Evaluation Scientist

COL Limited • Greater London

On-site
GBP 100,000 - 200,000
Market competitive salary
Equity options
Unlimited vacation
+2
AI/ML Engineer — RLHF & Model Evaluation (Remote)
AI/ML Engineer — RLHF & Model Evaluation (Remote)

Prolific • Greater London

On-site
GBP 80,000 - 100,000
AI Research Engineer - LLM Evaluation & Alignment (UK)
AI Research Engineer - LLM Evaluation & Alignment (UK)

Jobgether • United Kingdom

Hybrid
GBP 60,000 - 80,000
Open-Model AI Safety Lead: Red Team & Validation | Equity
Open-Model AI Safety Lead: Red Team & Validation | Equity

Reflection AI • Greater London

On-site
GBP 100,000 - 180,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Paid parental leave for all new parents
+2
AI/ML Engineer — RLHF & Model Evaluation (Remote)
AI/ML Engineer — RLHF & Model Evaluation (Remote)

Prolific • Sheffield

On-site
GBP 80,000 - 100,000
Staff AI Engineer: Post-Training, RL & Model Alignment
Staff AI Engineer: Post-Training, RL & Model Alignment

Reflection AI • Greater London

On-site
GBP 75,000 - 110,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+2
Staff Engineer, Safety for AI Agents
Staff Engineer, Safety for AI Agents

Cohere • Greater London

Hybrid
GBP 90,000 - 150,000
Open and inclusive culture
Work with cutting-edge AI research team
Weekly lunch stipend
+5
AI Model Evaluation Scientist - Pre-Deployment & Red-Teaming (Equity)
AI Model Evaluation Scientist - Pre-Deployment & Red-Teaming (Equity)

OpenDoorHub • Greater London

On-site
GBP 90,000 - 140,000
Equity
Remote AI Quality Analyst: Personalization & Evaluation
Remote AI Quality Analyst: Personalization & Evaluation

Turing • Greater London

On-site
GBP 45,000 - 60,000
AI Research Intern: Build & Validate AI Models
AI Research Intern: Build & Validate AI Models

Fastk • Birmingham

On-site
GBP 30,000 - 50,000