Senior Researcher, Evals

techire ai

San Francisco (CA)

Hybrid

USD 200,000 - 350,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Techire AI is seeking a Senior Research, Evals to rethink how model performance is measured across real-world interactions. You will develop evaluation frameworks that go beyond static benchmarks and build systems that integrate evaluations into model development.

The role is hands-on, bridging model development and evaluation, with responsibilities to design studies, pipelines and metrics for reasoning, memory, and interaction quality.

Qualifications

  • Experience building evaluation frameworks for generative models across text, audio or multimodal AI.
  • Strong technical and analytical skills.
  • Experience turning open-ended research questions into working evaluation systems.
  • A good understanding of metrics, experimentation and statistical analysis.
  • An interest in measuring subjective qualities such as naturalness, adaptability and interaction quality.

Responsibilities

  • Develop new evaluations for reasoning, memory, interaction and model behaviour.
  • Build evaluation pipelines with robust statistical analysis.
  • Create quantitative metrics for qualities that can be difficult to measure objectively.
  • Integrate evals directly into model training and research workflows.
  • Design user studies and behavioural experiments to understand real-world model performance.

Job description

How do you actually evaluate an AI model when accuracy on a benchmark only tells part of the story?

A well-funded AI company developing next-generation foundation models is looking for a Senior Research, Evals to rethink how model performance is measured across real-world interactions.

The role

You’ll work on evaluation frameworks that go beyond static benchmarks, looking at how models reason, remember, adapt and interact over time.

This is a hands‑on research role sitting between model development and evaluation. You’ll help define what should be measured, work out how to measure it, then build the systems that bring those evaluations directly into model development.

What you’ll do
  • Develop new evaluations for reasoning, memory, interaction and model behaviour
  • Build evaluation pipelines with robust statistical analysis
  • Create quantitative metrics for qualities that can be difficult to measure objectively
  • Integrate evals directly into model training and research workflows
  • Design user studies and behavioural experiments to understand real-world model performance
What you’ll bring
  • Experience building evaluation frameworks for generative models across text, audio or multimodal AI
  • Strong technical and analytical skills
  • Experience turning open-ended research questions into working evaluation systems
  • A good understanding of metrics, experimentation and statistical analysis
  • An interest in measuring subjective qualities such as naturalness, adaptability and interaction quality

Experience with alignment or model behaviour research would be useful, but isn’t essential.

You’ll work closely with researchers building new foundation models, helping understand whether changes are creating meaningful improvements rather than simply moving benchmark scores.

Base salary is $200k-$350k DOE + generous equity.

Based in San Francisco, New York or London, working hybrid.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Evaluation Scientist - Model Behavior & Metrics
Senior AI Evaluation Scientist - Model Behavior & Metrics

techire ai • San Francisco (CA)

Hybrid
USD 200,000 - 350,000
Evaluation Lead
Evaluation Lead

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Software Engineer - Research Platform, Consumer Devices
Senior Software Engineer - Research Platform, Consumer Devices

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 325,000
Software Engineer, Evaluation Platform / Infra
Software Engineer, Evaluation Platform / Infra

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 300,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research, Post-Training Evals
Research, Post-Training Evals

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

SupportFinity™ • New York (NY)

On-site
USD 140,000 - 200,000
Open and inclusive culture
Work on cutting-edge AI research
Weekly lunches and snacks
+6
Software Engineer, Evaluation Platform / Infra
Software Engineer, Evaluation Platform / Infra

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 300,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Evaluation Platform / Infra
Software Engineer, Evaluation Platform / Infra

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
AI Research Engineer
AI Research Engineer

TTN Talent • San Francisco (CA)

On-site
USD 200,000 - 400,000