Remote AI Benchmark Engineer & Researcher

Pathway

Palo Alto (CA)

Remote

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

An AI technology startup is seeking a Benchmarking Specialist in Palo Alto to design and execute ML evaluation benchmarks. You'll work closely with the R&D team to define data standards and maintain documentation. The ideal candidate has experience in ML/LLM evaluation and is fluent in English. This is a full-time position with remote work possibilities, targeting an immediate start date. Competitive compensation will be based on your profile and location.

Qualifications

  • Experience with ML/LLM evaluation, data science, or technical product roles.
  • Comfortable reading papers and translating technical details.
  • Fluent in English and respectful of others.

Responsibilities

  • Design and execute benchmarks to guide AI model evaluation.
  • Collaborate with R&D team to build evaluation infrastructure.
  • Maintain documentation for datasets and benchmarks.

Skills

ML/LLM evaluation
Data science
Technical product roles
Reading papers
High-quality data
Documentation

Job description

An AI technology startup is seeking a Benchmarking Specialist in Palo Alto to design and execute ML evaluation benchmarks. You'll work closely with the R&D team to define data standards and maintain documentation. The ideal candidate has experience in ML/LLM evaluation and is fluent in English. This is a full-time position with remote work possibilities, targeting an immediate start date. Competitive compensation will be based on your profile and location.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options
ML Safety & Benchmarking Research Engineer
ML Safety & Benchmarking Research Engineer

Apple Inc. • San Francisco (CA)

On-site
USD 181,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
AI Benchmarks & Evaluations Program Manager
AI Benchmarks & Evaluations Program Manager

Mercor • San Francisco (CA)

On-site
USD 120,000 - 200,000
Performance bonus structure
Equity grant
$15K relocation bonus
+7
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Benchmarking Research Engineer: Frontier Model Evaluations
Benchmarking Research Engineer: Frontier Model Evaluations

Refresh AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
ML Evaluation Engineer: Benchmark & Model Quality
ML Evaluation Engineer: Benchmark & Model Quality

Reducto • San Francisco (CA)

On-site
USD 100,000 - 130,000
Unlimited PTO
Daily free lunch
Reimbursed transportation
+3
Research Scientist (Remote/US/LATAM)
Research Scientist (Remote/US/LATAM)

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
Remote AI Evaluation Engineer — Build & Benchmark RL Tasks
Remote AI Evaluation Engineer — Build & Benchmark RL Tasks

Cross Border Talents • Michigan

On-site
USD 83,000 - 165,000
Fully remote
Flexible schedule
Remote CS Research Expert - Benchmark & ML Systems
Remote CS Research Expert - Benchmark & ML Systems

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 103,000
Remote work
Flexible schedule
Competitive hourly rate