Member of Technical Staff - ML Evaluation & Benchmarking

Arena Physica

New York (NY)

On-site

USD 175,000 - 250,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Premium health benefits
401(k)
Unlimited PTO
Lunch via Sharebite
Relocation support

Job summary

Arena Physica is hiring a Member of Technical Staff for Evaluations to lead the design, verification, and automation of evaluation pipelines for hardware-focused AI systems. You will craft analyses, calibrate scoring, and report findings to CTO and partners.

The role spans infrastructure, statistics, and domain knowledge, with collaboration across electrical engineering and product teams. This is a high-impact role based in the New York area.

Qualifications

  • Strong Python skills with experience building research or production infrastructure.
  • Experience designing evaluations, benchmarks, or metrics for ML systems.
  • Hands-on experience with LLM agents and analysis of agent trajectories.
  • Experience with experimental design and statistics.

Responsibilities

  • Own design and reliability of evaluation harnesses and automated pipelines.
  • Design ablations that isolate the effect of model changes on performance.
  • Analyze agent trajectories to identify performance clusters and root causes.
  • Collaborate to hand-craft synthetic datasets that test capabilities.
  • Design and calibrate LLM judges and scoring systems.

Skills

Python
Eval design
LLM agents
Experimental design
Communication
Automation bias
EE background
LLM judge

Job description

Arena Physica is hiring a Member of Technical Staff for Evaluations to lead the design, verification, and automation of evaluation pipelines for hardware-focused AI systems. You will craft analyses, calibrate scoring, and report findings to CTO and partners.

The role spans infrastructure, statistics, and domain knowledge, with collaboration across electrical engineering and product teams. This is a high-impact role based in the New York area.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Member of Technical Staff, AI Benchmarking
Senior Member of Technical Staff, AI Benchmarking

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 100,000 - 150,000
Competitive compensation including equity
Opportunity to shape AI development
Member of Technical Staff, Evaluations
Member of Technical Staff, Evaluations

Arena Physica • New York (NY)

On-site
USD 175,000 - 250,000
Premium health benefits
401(k)
Unlimited PTO
+2
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation
Staff Engineer - AI Evaluation & Research Execution
Staff Engineer - AI Evaluation & Research Execution

METR • Berkeley (CA)

Hybrid
USD 285,548 - 503,116
Catered lunch and dinner daily
In-office gym and shower
Unlimited PTO
+6
Member of Technical Staff - Data Science
Member of Technical Staff - Data Science

Arena Intelligence, Inc. • United States

On-site
USD 130,000 - 180,000
Competitive compensation & equity
Health and wellness benefits
Cutting-edge AI projects
+1
Staff Engineer, AI Evaluation Infrastructure
Staff Engineer, AI Evaluation Infrastructure

LinkedIn • Mountain View (CA)

Hybrid
USD 175,000 - 287,000
Evaluation Platform Engineer: Build Scalable ML Benchmarks
Evaluation Platform Engineer: Build Scalable ML Benchmarks

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Data Scientist, AI Evaluation & Insights
Data Scientist, AI Evaluation & Insights

Arena • San Francisco (CA)

On-site
Staff Engineer, Evaluation Infrastructure
Staff Engineer, Evaluation Infrastructure

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health & Wellness
Equity
Flexible time off