LLM Evaluation Engineering Lead

DeepRec.ai

Redwood City (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

High autonomy
Strong technical peers
Meaningful equity

Job summary

A leading AI technology firm in Redwood City is seeking an LLM Evaluations Engineering Lead. In this full-time position, you will be responsible for building evaluation systems for agentic LLMs, ensuring improved performance and reliability. Ideal candidates have strong software engineering skills and deep understanding of evaluation methodologies for machine learning. Join to work on impactful AI systems with the autonomy to shape their development.

Qualifications

  • Strong experience building evaluation systems for ML models, preferably LLMs.
  • Deep understanding of agentic failure modes such as tool misuse and hallucinated evidence.
  • Comfortable operating between research environments and production systems.
  • Turn eval failures into training signals (SFT / DPO / RL).
  • Experience building evaluation systems for ML models, with LLMs preferred.

Responsibilities

  • Build eval harnesses for agentic LLM systems, both offline and in-workflow.
  • Design evaluations for planning, execution, recovery, and safety.
  • Implement verifier-driven scoring and regression gates.
  • Turn evaluation failures into useful training signals.

Skills

Building evaluation systems for ML models
Python
Data pipelines
Test harnesses
Distributed execution
Reproducibility
Understanding of agentic failure modes
Reasoning about metrics

Job description

LLM Evaluations Engineering Lead – SF Bay Area (Onsite)

Full-time / Permanent

We’re partnering with a deep‑tech AI company building autonomous, agentic systems for complex physical and real‑world environments. The team operates at the edge of what’s possible today, designing AI systems that plan, act, recover, and improve over long horizons in high‑stakes settings.

They’re hiring an LLM Evaluations Engineering Lead to own the evaluation, verification, and regression layer for agentic LLM systems running end‑to‑end workflows. This is not a metrics‑only role; you’ll be building the guardrails that determine whether the system is actually getting better.

Why this role matters

As agentic LLM systems move into long‑horizon planning and execution, evals become the bottleneck.

  • Agents are actually improving
  • Changes introduce silent regressions
  • Uncertainty is shrinking or compounding
  • "success" reflects real‑world outcomes, not proxy metrics

Escalating incorrect evals means downstream systems fail. This role sits directly on that fault line.

What you’ll do
  • Build eval harnesses for agentic LLM systems (offline + in‑workflow)
  • Design evals for planning, execution, recovery, and safety
  • Implement verifier‑driven scoring and regression gates
  • Turn eval failures into training signals (SFT / DPO / RL)
What they’re looking for
  • Strong experience building evaluation systems for ML models (LLMs strongly preferred)
  • Excellent software engineering fundamentals:
    • Python
    • Data pipelines
    • Test harnesses
    • Distributed execution
    • Reproducibility
  • Deep understanding of agentic failure modes, including:
    • Tool misuse
    • Hallucinated evidence
    • Reward hacking
    • Brittle formatting and schema drift
  • Ability to reason about what to measure, not just how to measure it
  • Comfortable operating between research experimentation and production systems
Why join
  • Work on frontier agentic AI systems with real‑world consequences
  • Own a foundational layer that determines system reliability and progress
  • High autonomy, strong technical peers, and meaningful equity
  • Build evals that actually matter, not academic benchmarks
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Evaluation & Verification Lead
LLM Evaluation & Verification Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Senior LLM Evaluation Engineer
Senior LLM Evaluation Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
Senior AI Engineer
Senior AI Engineer

arosplatforms | AI Consulting & Services • North Township (IN)

On-site
USD 120,000 - 170,000
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
LLM Evaluation Engineer
LLM Evaluation Engineer

ThirdLaw | Runtime AI Safety • United States

Hybrid
USD 120,000 - 160,000
Market cash compensation
Above-market equity
Generous benefits