LLM Evaluation & Benchmarking Engineer

Capitolis

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques, along with the ability to adapt to new findings. Join a top AI startup committed to responsible development and user privacy in machine learning.

Qualifications

  • Knowledge of LLM evaluation and data curation methods.
  • Experience in designing LLM benchmarking methods.
  • Ability to shift focus as new findings emerge.

Responsibilities

  • Own LLM evaluation processes and methods for benchmarks.
  • Generate synthetic data and conduct benchmarking.
  • Deliver scalable and reproducible production code.
  • Develop new benchmarking methods for safety and helpfulness.
  • Co-author academic papers and presentations.

Skills

LLM evaluation
Data curation techniques
Designing benchmarking methods
Adaptability and flexibility

Job description

Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques, along with the ability to adapt to new findings. Join a top AI startup committed to responsible development and user privacy in machine learning.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer — LLM Evaluation
ML Engineer — LLM Evaluation

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
LLM Evaluation & Verification Lead
LLM Evaluation & Verification Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
LLM Safety Research Engineer
LLM Safety Research Engineer

Capitolis • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Benchmark Engineer — Elevate Model Evaluation
LLM Benchmark Engineer — Elevate Model Evaluation

Vals AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Relocation support
Health insurance
Lunch provided
+3
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
ML Research Engineer — LLM Safety
ML Research Engineer — LLM Safety

Capitolis • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Benchmarking Research Scientist
LLM Benchmarking Research Scientist

Anyone AI Inc. • Northern (KY)

Hybrid
USD 110,000 - 160,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
ML Engineer — LLM Privacy
ML Engineer — LLM Privacy

Capitolis • San Francisco (CA)

On-site
USD 150,000 - 210,000