ML Evaluation Architect for AI Benchmarks

Causal

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Causal is building a Large Physics foundation Model to predict and influence physical systems, starting with weather. We seek research engineers to design a central evaluation framework and reusable pipelines that span models and teams.

You will implement evaluation pipelines, benchmark suites, baselines, and visualization tools to turn results into shared, actionable understanding. Strong SWE and statistics are essential.

Qualifications

  • Strong software engineering skills and experience building data or evaluation pipelines at scale.
  • Experience turning research or model outputs into metrics, benchmarks, and visualizations that teams rely on.
  • Solid grasp of probability and statistics, with the judgment to design evaluations that measure what they claim to
  • Full-stack range: comfortable building both backend pipelines and the frontend tools people read results in.
  • Owns deliverables end-to-end, from collecting requirements to autonomously driving execution.

Responsibilities

  • Design and build a central, reusable evaluation framework that every model and every team runs through
  • Implement evaluation pipelines, benchmark suites, and baselines that make model quality measurable and comparable across efforts
  • Build the visualization and dashboard tools that turn raw results into shared, actionable understanding for the whole team
  • Establish sound statistical methodology for evaluation, so teams can distinguish real improvements from noise
  • Partner with research and domain teams to translate what "good" means in each domain into standardized, automated metrics

Skills

Software engineering
Data pipelines
Metrics & benchmarks
Statistics
Full-stack development

Job description

Causal is building a Large Physics foundation Model to predict and influence physical systems, starting with weather. We seek research engineers to design a central evaluation framework and reusable pipelines that span models and teams.

You will implement evaluation pipelines, benchmark suites, baselines, and visualization tools to turn results into shared, actionable understanding. Strong SWE and statistics are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, ML Evaluation & Metrics Platform
Staff Engineer, ML Evaluation & Metrics Platform

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Physicist for AI-Driven Causal Modeling
Staff Physicist for AI-Driven Causal Modeling

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Evaluation Engineer: Benchmark & Model Quality
ML Evaluation Engineer: Benchmark & Model Quality

Reducto • San Francisco (CA)

On-site
USD 100,000 - 130,000
Unlimited PTO
Daily free lunch
Reimbursed transportation
+3
Staff Engineer, AI Production & Customer Platforms
Staff Engineer, AI Production & Customer Platforms

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 190,000
ML Evaluation & Metrics Engineer — Hybrid Role
ML Evaluation & Metrics Engineer — Hybrid Role

WindBorne Systems • Atlanta (GA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental, health, and vision insurance
Unlimited PTO
+2
Multimodal Physics AI Researcher
Multimodal Physics AI Researcher

Causal • San Francisco (CA)

On-site
USD 180,000 - 280,000
Staff Engineer - AI Evaluation & Metrics Platform
Staff Engineer - AI Evaluation & Metrics Platform

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000