Evaluation Systems Engineer — ML Metrics & Pipelines

Mendable

San Francisco (CA)

Hybrid

USD 210,000 - 275,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive equity
PTO 15 days
Parental leave 12 weeks
Wellness stipend
Learning & Development
Team offsites
Sabbatical
401(k) plan
Pet insurance

Job summary

Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data - the boring-hard problem everyone building with LLMs eventually hits, solved.

You’ll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what “good” means and can measure it at scale, this is the role.

Qualifications

  • Experience: 4+ years in ML, research engineering, or data-heavy backend, with real evaluation work

Responsibilities

  • Design the metrics that define what 'good output' actually means across millions of sites, formats, and edge cases
  • Build the pipelines and harnesses that measure quality rigorously and at scale
  • Generate and curate the datasets that make evaluation trustworthy
  • Own the feedback loop from output quality back to model and product decisions
  • Turn 'did that work?' into an answer the whole team can act on

Skills

4+ years experience
ML
research engineering

Job description

Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data - the boring-hard problem everyone building with LLMs eventually hits, solved.

You’ll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what “good” means and can measure it at scale, this is the role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Evaluation Systems Engineer: Metrics, Datasets & Pipelines
Evaluation Systems Engineer: Metrics, Datasets & Pipelines

Visa Hunt • San Francisco (CA)

Hybrid
USD 210,000 - 275,000
Salary and equity
Generous PTO
Parental leave
+5
Eval Infrastructure Engineer — Define LLM Quality Metrics
Eval Infrastructure Engineer — Define LLM Quality Metrics

Firecrawl • San Francisco (CA)

Hybrid
USD 160,000 - 240,000
Unlimited PTO
12 weeks fully paid parental leave
Wellness stipend
+5
Customer-Facing Deployment Engineer
Customer-Facing Deployment Engineer

Mendable • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Medical, dental, and vision insurance
401(k) plan
Paid parental leave
+3
Staff Engineer, ML Evaluation & Metrics Platform
Staff Engineer, ML Evaluation & Metrics Platform

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Golden Data Pipelines
Senior Software Engineer - Golden Data Pipelines

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Research Engineer
Research Engineer

Visa Hunt • San Francisco (CA)

Hybrid
USD 210,000 - 275,000
Salary and equity
Generous PTO
Parental leave
+5
AI Engineer — Build Production-Grade LLM Pipelines
AI Engineer — Build Production-Grade LLM Pipelines

Royal Cyber • United States

Remote
USD 120,000 - 180,000
Lead ML Engineer: Scale Data Pipelines for LLMs
Lead ML Engineer: Scale Data Pipelines for LLMs

Cisco Systems, Inc. • Seattle (WA)

Hybrid
USD 180,000 - 260,000
Senior AI Evaluation Engineer: Pipelines & Metrics
Senior AI Evaluation Engineer: Pipelines & Metrics

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
ML Engineer: Annotation Pipelines & Evaluation
ML Engineer: Annotation Pipelines & Evaluation

RustLabs • Connecticut

On-site
USD 120,000 - 165,000