AI Model Evaluation Engineer — Benchmarking & Validation

SpreeAI

San Francisco (CA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A fast-growing AI company seeks a Software Engineer to focus on Model Evaluation & Benchmarking. This role involves building evaluation systems for multimodal AI, ensuring reliable performance. The ideal candidate will possess strong Python programming skills, familiarity with machine learning workflows, and experience in automation frameworks. Join us to redefine fashion and e-commerce with cutting-edge AI solutions. Passion for innovation in tech is a must. Enjoy a dynamic environment where creativity and technology drive real impact.

Qualifications

  • Degree in Computer Science, AI, Engineering, or comparable combination of education and practical experience.
  • Strong programming skills in Python.
  • Familiarity with object-oriented programming (C++, Java, Python, or similar).
  • Strong data structures and algorithms fundamentals.
  • Understanding of machine learning experimentation workflows.

Responsibilities

  • Build automated evaluation pipelines for multimodal AI models.
  • Benchmark diffusion models, vision systems, and generative workflows.
  • Validate model checkpoints and detect regressions across versions.
  • Develop evaluation metrics for realism, consistency, and performance.
  • Integrate evaluation tooling into CI/CD workflows.
  • Collaborate with ML researchers and infrastructure teams.
  • Analyze failure modes and propose evaluation strategies.

Skills

Programming skills in Python
Familiarity with object-oriented programming
Data structures and algorithms
Understanding of machine learning experimentation workflows

Education

Degree in Computer Science, AI, Engineering or equivalent

Tools

NumPy
Pandas

Job description

A fast-growing AI company seeks a Software Engineer to focus on Model Evaluation & Benchmarking. This role involves building evaluation systems for multimodal AI, ensuring reliable performance. The ideal candidate will possess strong Python programming skills, familiarity with machine learning workflows, and experience in automation frameworks. Join us to redefine fashion and e-commerce with cutting-edge AI solutions. Passion for innovation in tech is a must. Enjoy a dynamic environment where creativity and technology drive real impact.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Evaluation Engineer: Benchmark & Model Quality
ML Evaluation Engineer: Benchmark & Model Quality

Reducto • San Francisco (CA)

On-site
USD 100,000 - 130,000
Unlimited PTO
Daily free lunch
Reimbursed transportation
+3
AI Benchmarking Engineer — Evaluations & Failure Analysis
AI Benchmarking Engineer — Evaluations & Failure Analysis

Mercor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
+2
Applied AI Research Intern: Benchmark & Evaluation
Applied AI Research Intern: Benchmark & Evaluation

Labelbox • San Francisco (CA)

Hybrid
USD 35,000 - 45,000
Career advancement opportunities
Hybrid work model
Fast-paced environment
AI Evaluation Scientist: Psychometrics & Benchmarks
AI Evaluation Scientist: Psychometrics & Benchmarks

Experimentation Jobs • United States

On-site
USD 120,000 - 180,000
AI Model Behavior Engineer—Quality & Evaluation
AI Model Behavior Engineer—Quality & Evaluation

Notion • San Francisco (CA)

On-site
USD 98,000 - 140,000
AI Model Evaluation Specialist
AI Model Evaluation Specialist

BAM Ventures • New York (NY)

On-site
USD 100,000 - 130,000
Senior AIML Engineer — AI Model Evaluation & Benchmarking
Senior AIML Engineer — AI Model Evaluation & Benchmarking

Apple Inc. • Cupertino (CA)

On-site
USD 212,000 - 387,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
AI Systems Performance Engineer
AI Systems Performance Engineer

OpenAI • Los Angeles (CA)

Hybrid
USD 100,000 - 150,000
Relocation assistance
Hybrid work model
AI Risk & Fraud Evaluation Engineer
AI Risk & Fraud Evaluation Engineer

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+2