Founding AI Learning & Evaluation Scientist — Equity

Socket.dev

Beverly Hills (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Daily team dinner provided in-office

Job summary

Socket.dev in Beverly Hills is seeking a principal AI evaluation scientist to own the measurement layer under model training and learning science across StudyFetch and Honen. You will design evals, run benchmarks, and ensure safety and educational quality.

This founding-team role offers direct collaboration with decision-makers, end-to-end analytics ownership, and a culture of rigorous causality, transparent reporting, and learning-driven product decisions.

Qualifications

  • PhD in a quantitative field with 5+ years applying it to real products or research.
  • Proven track record in AI evaluation or measurement for products shipped to users.
  • Strong communication skills to present to engineers, founders, and researchers.

Responsibilities

  • Design evaluation frameworks and determine when a new model checkpoint is better.
  • Build and maintain an internal benchmark for multi-turn tutoring.
  • Link product data to model training using the Learn Engine.
  • Define instrumentation and telemetry to capture required events.
  • Publish parts of the benchmark to researchers outside the company.

Skills

Quantitative evaluation
Communication
Experiment design

Education

PhD in statistics / ML / CS / economics / physics

Tools

Python (Pandas/NumPy/SciPy/Jupyter)
SQL
MongoDB
PostgreSQL
GCP

Job description

Socket.dev in Beverly Hills is seeking a principal AI evaluation scientist to own the measurement layer under model training and learning science across StudyFetch and Honen. You will design evals, run benchmarks, and ensure safety and educational quality.

This founding-team role offers direct collaboration with decision-makers, end-to-end analytics ownership, and a culture of rigorous causality, transparent reporting, and learning-driven product decisions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding AI Learning & Evaluation Scientist
Founding AI Learning & Evaluation Scientist

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Socket.dev • Beverly Hills (CA)

On-site
USD 180,000 - 280,000
Daily team dinner provided in-office
Staff Engineer, Applied AI for Education
Staff Engineer, Applied AI for Education

StudyFetch • Beverly Hills (CA)

On-site
USD 170,000 - 270,000
100% employer-paid Medical, Dental, and Vision
75% dependent coverage
401(k) with employer matching
+1
Chief AI Evaluation & Research
Chief AI Evaluation & Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
Staff Engineer, AI Evaluation Infrastructure
Staff Engineer, AI Evaluation Infrastructure

LinkedIn • Mountain View (CA)

Hybrid
USD 175,000 - 287,000
Senior Staff AI Infrastructure Engineer, Model Evaluation
Senior Staff AI Infrastructure Engineer, Model Evaluation

LinkedIn • United States

Hybrid
USD 198,000 - 326,000
Evaluation Lead
Evaluation Lead

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Full-Stack Engineer for AI Evaluation Tools & Search (Equity)
Full-Stack Engineer for AI Evaluation Tools & Search (Equity)

Necessary Ventures • Sunnyvale (CA)

On-site
USD 120,000 - 180,000
Equity package
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000