LLM Evaluation & Benchmarking Engineer

Capitolis

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques, along with the ability to adapt to new findings. Join a top AI startup committed to responsible development and user privacy in machine learning.

Qualifications

  • Knowledge of LLM evaluation and data curation methods.
  • Experience in designing LLM benchmarking methods.
  • Ability to shift focus as new findings emerge.

Responsibilities

  • Own LLM evaluation processes and methods for benchmarks.
  • Generate synthetic data and conduct benchmarking.
  • Deliver scalable and reproducible production code.
  • Develop new benchmarking methods for safety and helpfulness.
  • Co-author academic papers and presentations.

Skills

LLM evaluation
Data curation techniques
Designing benchmarking methods
Adaptability and flexibility

Job description

Dynamo AI is seeking a candidate to lead LLM evaluation and benchmarking in San Francisco, California. You will generate high-quality data and develop innovative methods for assessing the safety and helpfulness of LLMs. The role requires domain knowledge in evaluation techniques, along with the ability to adapt to new findings. Join a top AI startup committed to responsible development and user privacy in machine learning.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer — LLM Evaluation
ML Engineer — LLM Evaluation

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
LLM Evaluation & Verification Lead
LLM Evaluation & Verification Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
LLM Evaluations Engineer — Benchmark Leaderboards
LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
LLM Benchmark Lead Research Scientist
LLM Benchmark Lead Research Scientist

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
LLM Benchmark Engineer — Lead Leaderboard Insights
LLM Benchmark Engineer — Lead Leaderboard Insights

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
ML Engineer: LLM Evaluation & Observability
ML Engineer: LLM Evaluation & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
Lead AI Engineer — LLM Evaluation & Production Optimizer
Lead AI Engineer — LLM Evaluation & Production Optimizer

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
Senior AI Engineer: LLM Evaluation, Production & Optimization
Senior AI Engineer: LLM Evaluation, Production & Optimization

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
AI Engineer: Build Production‑Ready LLM Features
AI Engineer: Build Production‑Ready LLM Features

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
US$1,000 per month food and commuting allowance
Laptop of choice
ESOP available