Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation

Palo Alto (CA)

Remote

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options

Job summary

An innovative AI startup is seeking a Benchmark Specialist to design and execute rigorous benchmarks and evaluate datasets for their post-transformer models. This role requires strong experience in ML/LLM evaluation and the ability to communicate technical specifications to both engineers and customers. The position is full-time and offers remote work flexibility. If you are passionate about high-quality data and groundbreaking research, this opportunity could be for you.

Qualifications

  • Experience with ML/LLM evaluation, data science, or technical product roles, ideally around benchmarks.
  • Comfortable reading papers, leaderboards, and Github repos.
  • Ability to communicate technical details in business terms.

Responsibilities

  • Design and execute rigorous benchmarks.
  • Evaluate candidate benchmarks for clarity and quality.
  • Organize benchmark results and maintain shared documentation.

Skills

ML/LLM evaluation
Data science
Technical product roles
High-quality data

Job description

An innovative AI startup is seeking a Benchmark Specialist to design and execute rigorous benchmarks and evaluate datasets for their post-transformer models. This role requires strong experience in ML/LLM evaluation and the ability to communicate technical specifications to both engineers and customers. The position is full-time and offers remote work flexibility. If you are passionate about high-quality data and groundbreaking research, this opportunity could be for you.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Benchmark Engineer & Researcher
Remote AI Benchmark Engineer & Researcher

Pathway • Palo Alto (CA)

Remote
USD 120,000 - 180,000
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
Applied AI Research Intern: Benchmark & Evaluation
Applied AI Research Intern: Benchmark & Evaluation

Labelbox • San Francisco (CA)

Hybrid
USD 35,000 - 45,000
Career advancement opportunities
Hybrid work model
Fast-paced environment
Remote Underwriting Expert — AI Benchmarking Projects
Remote Underwriting Expert — AI Benchmarking Projects

Mercor • San Francisco (CA)

On-site
USD 60,000 - 80,000
AI Benchmarking Engineer — Evaluations & Failure Analysis
AI Benchmarking Engineer — Evaluations & Failure Analysis

Mercor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
+2
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Independent SME AI Engineer – Domain Benchmarking
Independent SME AI Engineer – Domain Benchmarking

Lilt • United States

Remote
USD 90,000 - 140,000
Independent contractor
Competitive rates
Global community
+1
Remote CS Research Expert - Benchmark & ML Systems
Remote CS Research Expert - Benchmark & ML Systems

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 103,000
Remote work
Flexible schedule
Competitive hourly rate
Remote AI Benchmark QA Engineer
Remote AI Benchmark QA Engineer

24-Mag Llc • New York (NY)

Remote
USD 76,000 - 117,000