Staff Engineer, Evaluation Infrastructure

Simile

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health & Wellness
Equity
Flexible time off

Job summary

Simile is building the first AI simulation of society and is hiring for a Member of Technical Staff in Evaluations Engineering. You will create scalable systems to evaluate how well simulations match ground truth, with influence across data, backend services, automation, and internal tooling.

You will partner with Evals, Modeling, Product Engineering, and Data Operations. You will streamline evaluation runs, strengthen versioning and access controls, automate customer validations, and develop

Qualifications

  • Strong engineering fundamentals for production-quality software.
  • Experience building backend services, data pipelines, automation workflows, and relational data models.
  • Ability to take ambiguous projects from design through deployment and adoption.
  • Strong intuition for reliable evaluation infrastructure including versioning and reproducibility.
  • Familiarity with modern model-development and evaluation workflows.
  • Ability to build clear tools for researchers, engineers, and operators.
  • Proven track record of independently driving technical work across teams.

Responsibilities

  • Build evaluation execution infrastructure: pipelines and orchestration across datasets, models, populations, use cases.
  • Strengthen evaluation data systems: design relational schemas, versioning, provenance, permissions, quality controls.
  • Automate validation and data collection: streamline validations, survey deployment, response ingestion, ground truth integration.
  • Build human data workflows: create labeling and review tools for external experts and operators.
  • Develop evaluation tooling: interfaces to manage evals, compare models, investigate results, and identify regressions.

Skills

Engineering fundamentals
Data pipelines
End-to-end execution
Evaluation judgment
ML/LLM fluency
Product & user judgment
Ownership & communication

Job description

Simile is building the first AI simulation of society and is hiring for a Member of Technical Staff in Evaluations Engineering. You will create scalable systems to evaluate how well simulations match ground truth, with influence across data, backend services, automation, and internal tooling.

You will partner with Evals, Modeling, Product Engineering, and Data Operations. You will streamline evaluation runs, strengthen versioning and access controls, automate customer validations, and develop

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Evaluations Engineering - Member of Technical Staff
Evaluations Engineering - Member of Technical Staff

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health & Wellness
Equity
Flexible time off
Eval Infrastructure Tech Lead — Scalable AI Evaluation
Eval Infrastructure Tech Lead — Scalable AI Evaluation

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Staff Engineer — Evals & Post-Training Product
Staff Engineer — Evals & Post-Training Product

Fireworks AI • San Mateo (CA)

On-site
USD 120,000 - 160,000
Work with cutting-edge technology
Collaborative and inclusive environment
Opportunities for growth and learning
Staff Engineer - AI Evaluation & Research Execution
Staff Engineer - AI Evaluation & Research Execution

METR • Berkeley (CA)

Hybrid
USD 285,000 - 504,000
Catered lunch and dinner daily
In-office gym and shower
Unlimited PTO
+6
Staff Engineer, AI Evaluation & Metrics Platform
Staff Engineer, AI Evaluation & Metrics Platform

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 200,000 - 550,000
Equity in total compensation
401(k) with 6% matching
Health, dental, and vision insurance
+2
Staff Software Engineer, Data Eval & Simulation
Staff Software Engineer, Data Eval & Simulation

Perplexity AI • United States

On-site
USD 120,000 - 190,000
Eval Engineer for AI Model Evaluations
Eval Engineer for AI Model Evaluations

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Eval Infrastructure Tech Lead - Scale Safe AI Metrics
Eval Infrastructure Tech Lead - Scale Safe AI Metrics

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Equity option
Vacation and leave
Flexible hours
+1
Engineering Manager, AI Evaluation Systems
Engineering Manager, AI Evaluation Systems

Cursor • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff, Model Evaluations & Behavioral Metrics (Equity)
Staff, Model Evaluations & Behavioral Metrics (Equity)

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health & Wellness benefits
Flexible time off
Equity grants