AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

A cutting-edge AI technology firm in San Francisco is seeking an Evaluation Lead to drive the assessment of AI model performance. You will design evaluation methodologies, automate evaluation processes, and oversee various evaluation strategies. The ideal candidate has extensive experience in AI model evaluation and is proficient in Python. This high-impact role demands strong collaboration skills and a startup-ready mindset to thrive in a fast-paced environment.

Qualifications

  • Extensive expertise in evaluating AI and machine learning models, ideally in physical AI.
  • Experience in designing, implementing, and refining evaluation metrics.
  • Deep understanding of machine learning, AI, and generative models.

Responsibilities

  • Design and implement evaluation methodologies and benchmarks for model effectiveness.
  • Build and oversee pipelines and tools that automate model evaluation.
  • Develop strategies for evaluating physical AI models across various use cases.

Skills

Evaluating AI and machine learning models
Designing evaluation metrics
Understanding of machine learning
Python programming
Building scalable data pipelines
Strong communication skills
Collaboration with teams

Job description

A cutting-edge AI technology firm in San Francisco is seeking an Evaluation Lead to drive the assessment of AI model performance. You will design evaluation methodologies, automate evaluation processes, and oversee various evaluation strategies. The ideal candidate has extensive experience in AI model evaluation and is proficient in Python. This high-impact role demands strong collaboration skills and a startup-ready mindset to thrive in a fast-paced environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmarks & Evaluations Program Manager
AI Benchmarks & Evaluations Program Manager

Mercor • San Francisco (CA)

On-site
USD 120,000 - 200,000
Performance bonus structure
Equity grant
$15K relocation bonus
+7
AI Risk & Fraud Evaluation Engineer
AI Risk & Fraud Evaluation Engineer

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+2
Head of AI Evaluation & Benchmarks
Head of AI Evaluation & Benchmarks

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
Chief AI Evaluation & Research
Chief AI Evaluation & Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation
Remote AI Evaluation Engineer: Safeguard & Scale
Remote AI Evaluation Engineer: Safeguard & Scale

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Senior AI Evaluation Engineer — Metrics & Data Pipelines
Senior AI Evaluation Engineer — Metrics & Data Pipelines

Sentry • San Francisco (CA)

Hybrid
USD 240,000 - 280,000
Equity grants
Paid time off
Group health insurance coverage
AI Evaluation Engineer - Build Eval Suites & Quality Gates
AI Evaluation Engineer - Build Eval Suites & Quality Gates

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Infrastructure Performance Modeling Lead
AI Infrastructure Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
AI Evaluation Engineer (Contract) — Python QA & Testing
AI Evaluation Engineer (Contract) — Python QA & Testing

Mindrift • Colorado

Remote
USD 80,000 - 100,000