LLM Evaluation & Verification Lead

DeepRec.ai

Redwood City (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

High autonomy
Strong technical peers
Meaningful equity

Job summary

A leading AI technology firm in Redwood City is seeking an LLM Evaluations Engineering Lead. In this full-time position, you will be responsible for building evaluation systems for agentic LLMs, ensuring improved performance and reliability. Ideal candidates have strong software engineering skills and deep understanding of evaluation methodologies for machine learning. Join to work on impactful AI systems with the autonomy to shape their development.

Qualifications

  • Strong experience building evaluation systems for ML models, preferably LLMs.
  • Deep understanding of agentic failure modes such as tool misuse and hallucinated evidence.
  • Comfortable operating between research environments and production systems.

Responsibilities

  • Build eval harnesses for agentic LLM systems, both offline and in-workflow.
  • Design evaluations for planning, execution, recovery, and safety.
  • Implement verifier-driven scoring and regression gates.
  • Turn evaluation failures into useful training signals.

Skills

Building evaluation systems for ML models
Python
Data pipelines
Test harnesses
Distributed execution
Reproducibility
Understanding of agentic failure modes
Reasoning about metrics

Job description

A leading AI technology firm in Redwood City is seeking an LLM Evaluations Engineering Lead. In this full-time position, you will be responsible for building evaluation systems for agentic LLMs, ensuring improved performance and reliability. Ideal candidates have strong software engineering skills and deep understanding of evaluation methodologies for machine learning. Join to work on impactful AI systems with the autonomy to shape their development.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
LLM Evaluation & Benchmarking Engineer
LLM Evaluation & Benchmarking Engineer

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
LLM Evaluation Scenario Architect
LLM Evaluation Scenario Architect

Mindrift • Mississippi

Remote
USD 10,000 - 60,000
Lead AI Engineer — LLM Evaluation & Production Optimizer
Lead AI Engineer — LLM Evaluation & Production Optimizer

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Remote Applied AI Research Scientist (LLM & Evaluation)
Remote Applied AI Research Scientist (LLM & Evaluation)

Rex.zone • United States

Remote
USD 80,000 - 100,000
Senior LLM Engineer - End-to-End Production AI
Senior LLM Engineer - End-to-End Production AI

Albiware Inc. • Chicago (IL)

On-site
USD 140,000 - 210,000
Competitive salary
Generous PTO
Medical, dental, and vision coverage
+2
Senior LLM Engineer: Build Production-Grade AI Systems
Senior LLM Engineer: Build Production-Grade AI Systems

Albi • Chicago (IL)

On-site
USD 150,000 - 210,000
Competitive salary
Generous PTO
Medical, dental, and vision coverage
+2
ML Engineer: LLMs, VLMs & Reasoning AI | Equity
ML Engineer: LLMs, VLMs & Reasoning AI | Equity

Tensor • San Jose (CA)

On-site
USD 75,000 - 300,000
Competitive compensation package
Participation in discretionary equity incentive plan
Access to comprehensive benefits program
Senior AI Engineer - Production LLM EvalOps
Senior AI Engineer - Production LLM EvalOps

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000