AI Evaluation Infrastructure Engineer

Block

Sacramento (CA)

Remote

USD 264,000 - 395,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote work
Medical insurance
Flexible time off
Retirement savings plans
Family planning

Job summary

Block is hiring an engineer to build infrastructure and tooling for high-quality AI evaluation at scale. You will create systems to compare candidate versions, measure accuracy, and ensure offline and online results align.

The team collaborates with product, data, and ML groups to deliver fast, trustworthy evaluation workflows. This role emphasizes reliable, observable systems, rapid iteration, and measurable improvements in evaluation quality that inform product decisions across Block's AI

Qualifications

  • Experience building production platforms or infrastructure, including distributed batch execution, data pipelines, or systems that process production logs.
  • Strong statistical literacy, including confidence intervals, variance, power, and multiple comparisons.
  • Judgment to identify when a result is meaningful and when it is noise.
  • Experience evaluating LLM or ML systems, or deep systems engineering with a strong interest in AI evaluation.
  • Product instinct for internal tools. You understand that leaderboards, annotation tools, and workflows only matter if teams actually use them.
  • A bias toward building reliable, observable systems that other engineers can trust.
  • Strong collaboration skills across ambiguous product, data, and engineering problems.
  • In your first year: product teams can stand up credible evals for new AI surfaces in days, and CI gates run quickly.

Responsibilities

  • Build an execution engine that can score candidate versions against task sets in minutes, not hours.
  • Create task set tooling that samples from production logs and validates tasks before they are admitted into an evaluation set.
  • Build grader infrastructure across ground truth checks, rubrics, and LLM-as-judge approaches.
  • Develop tooling that helps human reviewers calibrate judges, measure judge-to-human agreement, and monitor drift over time.
  • Build leaderboards and reporting systems that include sample size, confidence intervals, and run-to-run variance.

Skills

Distributed batch execution
Data pipelines
Production logs
LLM/ML evaluation
Observability
Cross-functional collaboration

Tools

Distributed systems tooling

Job description

It all started with an idea at Block in 2013. Initially built to take the pain out of peer-to-peer payments, Cash App has gone from a simple product with a single purpose to a dynamic ecosystem, developing unique financial products, including Afterpay/Clearpay, to provide a better way to send, spend, invest, borrow and save to our 50+ million monthly active customers. We want to redefine the world’s relationship with money to make it more relatable, instantly available, and universally accessible.Today, Cash App has thousands of employees working globally across office and remote locations, with a culture geared toward innovation, collaboration and impact. We’ve been a distributed team since day one, and many of our roles can be done remotely from the countries where Cash App operates. No matter the location, we tailor our experience to ensure our employees are creative, productive, and happy.

The Role

We build AI products, and the quality of our evaluations sets the ceiling for how good those products can be. The speed of our evaluations determines how quickly we can improve them.We are looking for an engineer to build the infrastructure and tooling that make high-quality AI evaluation possible at Block's scale. You will help teams understand whether a model or product change is actually better, whether a result is statistically meaningful, and whether offline evaluation is predicting what happens with real users.Our evaluation approach combines offline evals that encode our definition of a good response, online evals that show how people actually respond, and a feedback loop that keeps the two converging. Your work will turn that approach into systems that product teams can use quickly, reliably, and with confidence.This is a high-impact, early-stage area with broad surface area. You will help decide what to build first, then build the platform that helps teams ship better AI products faster.

You Will
  • Build an execution engine that can score candidate versions against task sets in minutes, not hours.
  • Create task set tooling that samples from production logs and validates tasks before they are admitted into an evaluation set.
  • Build grader infrastructure across ground truth checks, rubrics, and LLM-as-judge approaches.
  • Develop tooling that helps human reviewers calibrate judges, measure judge-to-human agreement, and monitor drift over time.
  • Build leaderboards and reporting systems that include sample size, confidence intervals, and run-to-run variance, so teams can distinguish real improvements from noise.
  • Support in-product side-by-side serving, feedback capture, and implicit signal extraction from real conversations.
  • Build the loop that compares offline scores with online outcomes, identifies eval sets that have stopped predicting reality, and helps teams improve them.
  • Partner with product, engineering, data, and ML teams to make evaluation workflows fast enough and trustworthy enough to become part of everyday development.
You Have
  • Experience building production platforms or infrastructure, including distributed batch execution, data pipelines, or systems that process production logs.
  • Strong statistical literacy, including comfort with confidence intervals, variance, power, and multiple comparisons.
  • The judgment to identify when a result is meaningful and when it is noise.
  • Experience evaluating LLM or ML systems, or deep systems engineering experience with a strong interest in AI evaluation.
  • Product instinct for internal tools. You understand that leaderboards, annotation tools, and workflows only matter if teams actually use them.
  • A bias toward building reliable, observable systems that other engineers can trust.
  • Strong collaboration skills and the ability to work across ambiguous product, data, and engineering problems.
  • Success in your first year
    • Product teams can stand up credible evals for new AI surfaces in days.
    • Eval results gate CI and run quickly enough that engineers do not route around them.
    • Judge-to-human agreement is measured, published, and monitored for drift.
    • Launch decisions are not made on results that are within statistical noise.
    • At least one case is documented where online reality disagreed with offline scores, and the evaluation was improved as a result.
Why this matters

AI product development moves quickly, but speed only helps when teams can trust the signal they are using to make decisions. This role will build the systems that make those signals faster, more accurate, and more actionable. Your work will directly influence how Block evaluates, improves, and ships AI products.Block takes a market-based approach to pay, and pay may vary depending on your location. U.S. locations are categorized into one of four zones based on a cost of labor index for that geographic area. The successful candidate’s starting pay will be determined based on job-related skills, experience, qualifications, work location, and market conditions. These ranges may be modified in the future.To find a location’s zone designation, please refer to this resource. If a location of interest is not listed, please speak with a recruiter for additional information.

  • Zone A:$263,600—$395,400 USD
  • Zone B:$263,600—$395,400 USD
  • Zone C:$263,600—$395,400 USD
  • Zone D:$263,600—$395,400 USD
Benefits

Every benefit we offer is designed with one goal: empowering you to do the best work of your career while building the life you want. Remote work, medical insurance, flexible time off, retirement savings plans, and modern family planning are just some of our offering.

Check out our other benefits at Block.Block, Inc. builds technology to increase access to the global economy. Each of our brands unlocks different aspects of the economy for more people.

  • Square makes commerce and financial services accessible to sellers.
  • Cash App is the easy way to spend, send, and store money.
  • Afterpay is transforming the way customers manage their spending over time.
  • TIDAL is a music platform that empowers artists to thrive as entrepreneurs.
  • Bitkey is a simple self-custody wallet built for bitcoin.
  • Proto is a suite of bitcoin mining products and services. Together, we’re helping build a financial system that is open to everyone.

Privacy Policy

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Infrastructure Engineer
AI Evaluation Infrastructure Engineer

Block • San Francisco (CA)

Remote
USD 264,000 - 395,000
AI Evaluation Infrastructure Engineer Bay Area, CA, US
AI Evaluation Infrastructure Engineer Bay Area, CA, US

Block, Inc. • Northern (KY)

Hybrid
USD 264,000 - 395,000
Remote work
Medical insurance
Flexible time off
+2
Sr. People Tech Analyst
Sr. People Tech Analyst

Block, Inc. • San Francisco (CA)

On-site
USD 172,000 - 258,000
Remote work options
Medical insurance
Flexible time off
+2
Staff Applied Machine Learning Engineer - Intelligent Data, Signals & Systems Bay Area, CA, US
Staff Applied Machine Learning Engineer - Intelligent Data, Signals & Systems Bay Area, CA, US

Block, Inc. • Northern (KY)

On-site
USD 277,000 - 415,000
Remote work option
Medical insurance
Flexible time off
+2
Senior Analytics Engineer, AI & DX Analytics
Senior Analytics Engineer, AI & DX Analytics

Block, Inc. • New York (NY)

Remote
USD 139,000 - 233,000
Remote work
Medical insurance
Flexible time off
+2
Staff Software Engineer, Go-to-Market Systems & AI
Staff Software Engineer, Go-to-Market Systems & AI

Square • Oakland (CA)

On-site
USD 264,000 - 395,000
Sr. People Tech Analyst
Sr. People Tech Analyst

Block • Chicago (IL)

On-site
USD 147,000 - 245,000
Remote work
Medical insurance
Flexible time off
+1
Senior Analytics Engineer, AI & DX Analytics Bay Area, CA, US
Senior Analytics Engineer, AI & DX Analytics Bay Area, CA, US

Block, Inc. • California (MO)

Hybrid
USD 139,000 - 233,000
Remote work
Medical insurance
Flexible time off
+2
Senior Data Scientist, First Line Risk Bay Area, CA, US
Senior Data Scientist, First Line Risk Bay Area, CA, US

Block, Inc. • Northern (KY)

Remote
USD 168,000 - 297,000
Software Engineer, Finance Applications
Software Engineer, Finance Applications

Block • San Francisco (CA)

On-site
USD 218,000 - 327,000