AI Evaluation Scientist — Real-World ML Research

Arena Intelligence, Inc.

San Francisco (CA)

On-site

USD 170,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity
Health benefits
Small, mission-driven team
Impactful projects

Job summary

Arena Intelligence, Inc. in San Francisco seeks machine learning scientists to design and analyze experiments that reveal what makes AI models useful, trustworthy and capable through human preference signals. You will work with engineers, product teams, and researchers to develop methods for comparing models and disentangling performance factors.

This role is deeply interdisciplinary. You’ll contribute to open-ended questions, rigorous evaluation, and research with real-world impact while

Qualifications

  • PhD or equivalent research experience in ML, NLP, or Statistics.
  • Strong ability to design and analyze experiments with statistical rigor.
  • Experience publishing research or contributing to ML/NLP evaluation efforts.

Responsibilities

  • Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions.
  • Develop new metrics, methodologies, and evaluation protocols beyond traditional benchmarks.
  • Analyze large-scale human voting data to uncover insights into model performance and user preferences.
  • Collaborate with engineers to implement research findings into production systems.
  • Prototype and test research ideas quickly with rigor and speed.
  • Author internal reports and external publications contributing to ML research community.
  • Partner with model providers to shape evaluation questions and testing approaches.
  • Contribute to scientific integrity and transparency of Arena Intelligence leaderboard and tools.

Skills

Hands-on model training
ML & statistics foundations
Experiment design & analysis
Collaborative mindset

Education

PhD or equivalent

Tools

PyTorch
JAX
TensorFlow

Job description

Arena Intelligence, Inc. in San Francisco seeks machine learning scientists to design and analyze experiments that reveal what makes AI models useful, trustworthy and capable through human preference signals. You will work with engineers, product teams, and researchers to develop methods for comparing models and disentangling performance factors.

This role is deeply interdisciplinary. You’ll contribute to open-ended questions, rigorous evaluation, and research with real-world impact while

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied AI Engineer — Build & Eval Real‑World AI
Applied AI Engineer — Build & Eval Real‑World AI

Precision Labs • Northern (KY)

Hybrid
USD 110,000 - 170,000
Competitive compensation & equity
Health & wellness benefits
Work on cutting‑edge AI
+1
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning Scientist - ML Research
Machine Learning Scientist - ML Research

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Health benefits
+2
Client-Facing AI Engineer: Model Integration & Eval
Client-Facing AI Engineer: Model Integration & Eval

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Competitive compensation & equity
Health benefits
Cutting-edge AI projects
+1
Open-Source ML Scientist & Research Lead
Open-Source ML Scientist & Research Lead

Arena • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Research Engineer: Fine-Tuning & Evaluation Systems
ML Research Engineer: Fine-Tuning & Evaluation Systems

HonestAI • San Francisco (CA)

On-site
USD 140,000 - 190,000
Machine Learning Scientist - Open Source Lead
Machine Learning Scientist - Open Source Lead

Arena • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation & equity
Health and wellness benefits
Small, mission-driven team
+1
Applied AI Engineer
Applied AI Engineer

Precision Labs • Northern (KY)

Hybrid
USD 110,000 - 170,000
Competitive compensation & equity
Health & wellness benefits
Work on cutting‑edge AI
+1
Applied AI Engineer
Applied AI Engineer

Arena • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation
Equity
Health benefits
+1
Applied AI Engineer
Applied AI Engineer

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Competitive compensation & equity
Health benefits
Cutting-edge AI projects
+1