AI Evaluation Platform Engineer

Block

San Francisco (CA)

Remote

USD 264,000 - 395,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Block in San Francisco is hiring an engineer to build the infrastructure for high‑quality AI evaluation at scale. You will help teams determine if model changes improve outcomes, and if offline scores predict online behavior.

You will build systems that run fast, measure confidence, and ensure trustworthy evaluation across product, data, and engineering teams. The role emphasizes reliability, observability, and collaboration to ship better AI products.

Qualifications

  • Experience building production platforms or infrastructure, including distributed batch execution, data pipelines, or systems that process production logs.
  • Strong statistical literacy, including confidence intervals, variance, power, and multiple comparisons.
  • The judgment to identify when a result is meaningful and when it is noise.
  • Experience evaluating LLM or ML systems, or deep systems engineering with a strong interest in AI evaluation.
  • Product instinct for internal tools; leaderboards, annotation tools, and workflows must be used by teams.
  • A bias toward building reliable, observable systems that other engineers can trust.
  • Strong collaboration skills across ambiguous product, data, and engineering problems.

Responsibilities

  • Build an execution engine that can score candidate versions against task sets in minutes, not hours.
  • Create task set tooling that samples from production logs and validates tasks before admission into an evaluation set.
  • Build grader infrastructure across ground truth checks, rubrics, and LLM-as-judge approaches.
  • Develop tooling that helps human reviewers calibrate judges, measure judge-to-human agreement, and monitor drift over time.
  • Build leaderboards and reporting systems with sample size, confidence intervals, and run-to-run variance.
  • Support in-product side-by-side serving, feedback capture, and implicit signal extraction from real conversations.
  • Build the loop that compares offline scores with online outcomes and improves eval sets.
  • Partner with product, engineering, data, and ML teams to make evaluation workflows fast and trustworthy.

Skills

Production platforms
Statistics
Signal interpretation
LLM/ML evaluation
Internal tools mindset
Reliable systems
Cross-team collaboration

Job description

Block in San Francisco is hiring an engineer to build the infrastructure for high‑quality AI evaluation at scale. You will help teams determine if model changes improve outcomes, and if offline scores predict online behavior.

You will build systems that run fast, measure confidence, and ensure trustworthy evaluation across product, data, and engineering teams. The role emphasizes reliability, observability, and collaboration to ship better AI products.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Platform Engineer (Remote)
AI Evaluation Platform Engineer (Remote)

Block • Sacramento (CA)

Remote
USD 264,000 - 395,000
Remote work
Medical insurance
Flexible time off
+2
Remote AI Evaluation Infrastructure Engineer
Remote AI Evaluation Infrastructure Engineer

Block, Inc. • Northern (KY)

Hybrid
USD 264,000 - 395,000
Remote work
Medical insurance
Flexible time off
+2
Senior Platform Engineer — Go, APIs & AI Evaluation
Senior Platform Engineer — Go, APIs & AI Evaluation

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Health benefits
Competitive salary
+1
AI Evaluation Platform Engineer
AI Evaluation Platform Engineer

Ellipsis Health Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 210,000
401(k) matching
Health insurance
Flex time off
On-Device AI Evaluation Architect
On-Device AI Evaluation Architect

Rnb Consultancy • San Francisco (CA)

On-site
USD 200,000 - 250,000
Equity
Relocation support
Visa/immigration assistance
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Evaluation Platform Engineer: Build Scalable ML Benchmarks
Evaluation Platform Engineer: Build Scalable ML Benchmarks

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
AI Model Evaluation Engineer – Build Verifications
AI Model Evaluation Engineer – Build Verifications

Zof AI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 190,000
MacBook Pro
Premium AI development tools
OpenAI Codex Max
Software Engineer, Platform
Software Engineer, Platform

People Culture Talent • San Francisco (CA)

On-site
USD 200,000 - 350,000
Comprehensive health, dental, vision
Visa sponsorship available
Equity opportunity
Delivery Engineer: AI Safety & Enterprise Evaluations
Delivery Engineer: AI Safety & Enterprise Evaluations

Artificial Intelligence Underwriting Company • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive salary
Equity
Relocation to San Francisco
+1