AI Benchmarking Engineer — Evaluations & Failure Analysis

Mercor

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
Free Equinox membership
Health insurance

Job summary

A cutting-edge AI firm in San Francisco is seeking a Research Engineer to develop evaluation systems and benchmarking pipelines for language models. Candidates should have a strong background in applied research, coding skills, and familiarity with ML models. You will work collaboratively with cross-functional teams and operate in a fast-paced environment, conducting failure analyses and contributing to the improvement of AI tools. The role requires working onsite five days a week, emphasizing strong ownership and high-intensity collaboration.

Qualifications

  • Experience in model evaluation, benchmarking, or failure analysis is expected.
  • Hands-on coding experience with machine learning models and evaluation code is crucial.
  • Comfortable with cloud services for deploying and storing evaluation results.

Responsibilities

  • Design, implement, and maintain benchmarking systems.
  • Build LLM evaluation systems with scoring and tracking capabilities.
  • Conduct systematic failure analyses to refine training processes.
  • Collaborate across teams to align benchmarks with training objectives.

Skills

Applied research background
Strong coding skills
Data structures and algorithms
Comfort with APIs

Tools

ML models
SQL/NoSQL
Cloud platforms

Job description

A cutting-edge AI firm in San Francisco is seeking a Research Engineer to develop evaluation systems and benchmarking pipelines for language models. Candidates should have a strong background in applied research, coding skills, and familiarity with ML models. You will work collaboratively with cross-functional teams and operate in a fast-paced environment, conducting failure analyses and contributing to the improvement of AI tools. The role requires working onsite five days a week, emphasizing strong ownership and high-intensity collaboration.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmarking Engineer — Evaluation & Failure Analysis
AI Benchmarking Engineer — Evaluation & Failure Analysis

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6
AI Benchmarking & Evaluation Engineer
AI Benchmarking & Evaluation Engineer

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Remote AI Benchmark Engineer & Researcher
Remote AI Benchmark Engineer & Researcher

Pathway • Palo Alto (CA)

Remote
USD 120,000 - 180,000
Benchmarking Research Engineer: Frontier Model Evaluations
Benchmarking Research Engineer: Frontier Model Evaluations

Refresh AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Benchmarks & Evaluations Program Manager
AI Benchmarks & Evaluations Program Manager

Mercor • San Francisco (CA)

On-site
USD 120,000 - 200,000
Performance bonus structure
Equity grant
$15K relocation bonus
+7
AI Risk & Fraud Evaluation Engineer
AI Risk & Fraud Evaluation Engineer

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+2
AI Model Evaluation Engineer — Benchmarking & Validation
AI Model Evaluation Engineer — Benchmarking & Validation

SpreeAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
ML Safety & Benchmarking Research Engineer
ML Safety & Benchmarking Research Engineer

Apple Inc. • San Francisco (CA)

On-site
USD 181,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation