Remote AI Benchmark & Datasets Engineer

Ignite Next GmbH

Palo Alto, Northern (CA, KY)

Hybrid

USD 120,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote work
Office visits Palo Alto, Paris, Wroclw

Job summary

Pathway, headquartered in Palo Alto, is seeking a benchmarking-focused ML researcher. You will design and run rigorous benchmarks for post-transformer models, curate datasets, and translate research into repeatable specs for engineering and GTM teams.

You will collaborate with our R&D and product groups, track model performance, and help shape demonstrations and customer-ready results that showcase our evaluation framework. Fluency in English and strong communication are essential.

Qualifications

  • Publish at NeurIPS, ICLR or ICML as lead author or with significant contributions.
  • Contributed to an LLM training effort that gained notable attention.
  • Spent at least 6 months in a leading ML research center (e.g., Google Brain, DeepMind, Apple, Meta, Anthropic, Nvidia).
  • ICPC World Finalist or IOI/IMO/IPhO medalist in high school.
  • Experience with ML/LLM evaluation, data science or technical product roles focused on benchmarks.

Responsibilities

  • Design and execute benchmarks and define dataset standards.
  • Collaborate with R&D to build evaluation infrastructure for post‑transformer models.
  • Identify, curate benchmarks and develop repeatable benchmark specs.
  • Maintain documentation and track benchmark results and leaderboards.
  • Contribute to demos and public proof points based on benchmark outcomes.

Skills

Fluent English
Benchmarking
LLM Evaluation
Reading Papers
Interdisciplinary Communication

Job description

Pathway, headquartered in Palo Alto, is seeking a benchmarking-focused ML researcher. You will design and run rigorous benchmarks for post-transformer models, curate datasets, and translate research into repeatable specs for engineering and GTM teams.

You will collaborate with our R&D and product groups, track model performance, and help shape demonstrations and customer-ready results that showcase our evaluation framework. Fluency in English and strong communication are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options
AI Benchmark & Datasets Engineer / Researcher
AI Benchmark & Datasets Engineer / Researcher

Ignite Next GmbH • Palo Alto (CA), Northern (KY)

Hybrid
USD 120,000 - 190,000
Remote work
Office visits Palo Alto, Paris, Wroclw
AI Benchmark & Datasets Engineer / Researcher
AI Benchmark & Datasets Engineer / Researcher

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 150,000 - 210,000
Intellectually stimulating work environment
Work with a pioneering AI startup
Flexible remote work options
Remote ML Engineer: Build Benchmarks & Production Pipelines
Remote ML Engineer: Build Benchmarks & Production Pipelines

raydar • Northern (KY)

Hybrid
USD 170,000 - 270,000
Competitive equity
Remote AI Inference Benchmark Engineer
Remote AI Inference Benchmark Engineer

Silicon Data • United States

Remote
USD 140,000 - 200,000
AI Benchmarking & Performance Architect
AI Benchmarking & Performance Architect

CoreWeave • Sunnyvale (CA)

On-site
USD 206,000 - 333,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+2
Remote AI Benchmarking Engineer: Performance Economics
Remote AI Benchmarking Engineer: Performance Economics

Akka • United States

On-site
USD 140,000 - 230,000
Competitive salary
Health benefits
Professional development
+3
Remote Forward-Deployed ML Engineer — Benchmark & Eval
Remote Forward-Deployed ML Engineer — Benchmark & Eval

Advatix • Northern (KY)

Hybrid
USD 170,000 - 270,000
Senior AI Product Manager — Coding Benchmarks & Data
Senior AI Product Manager — Coding Benchmarks & Data

Scale AI, Inc. • New York (NY)

On-site
USD 206,000 - 257,000
Health insurance
Dental coverage
Vision coverage
+4
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation