Applied Research — Evaluations & Data

Human Intuition Inc.

New York (NY)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Human Intuition Inc. is building the autonomous company and seeks a candidate who can turn workflows, decision histories, and outcomes into datasets and evaluations that distinguish plausible actions from correct ones.

You will design tasks, evaluation environments, and verifiers, build data pipelines with provenance, and create repeatable reports to compare quality, reliability, latency, and cost.

Qualifications

  • Strong Python and practical experience building data or ML systems.
  • Evidence of designing useful evaluations, experiments, or datasets for models or complex software.
  • Sound statistical judgment and the ability to explain what a metric establishes and what it misses.
  • Care with data quality, reproducibility, and the handling of business information.
  • The ability to learn an unfamiliar domain and communicate clearly with both technical and operational colleagues.

Responsibilities

  • Work with domain experts to identify the decisions, constraints, and outcomes that matter in a workflow.
  • Design tasks, evaluation environments, rubrics, and verifiers that test complete workflows, including ambiguous inputs and failure recovery.
  • Build data pipelines that preserve the relationship between context, decisions, actions, and outcomes, with clear provenance and appropriate access controls.
  • Create evaluation splits and review processes that limit leakage, expose distribution shifts, and catch regressions.
  • Study agent traces to identify recurring failures and turn the findings into better training data, tools, and reward signals.
  • Build repeatable evaluation reports that let researchers and engineers compare quality, reliability, latency, and cost.

Skills

Python
Data/ML systems
Evaluation design
Statistical judgment
Data quality & reproducibility
Domain learning & communication

Job description

Building the autonomous company

Human Intuition is building the autonomous company. Businesses run on accumulated judgment: how to interpret a situation, choose an action, and learn from its consequences. Much of that knowledge lives in people, even when the decisions they make leave traces in software.

We are working to make that judgment learnable. A business has defined systems, tools, permissions, histories, and objectives. Those boundaries create an opportunity to build agents that learn from how work is done, act within clear constraints, and improve through feedback. Our ambition is to turn the knowledge inside institutions into software that compounds.

The role

Build the evidence that tells us whether an agent is ready to do useful work. You will turn business workflows, decision histories, and operational outcomes into datasets and evaluations that distinguish a plausible response from a correct action. Your work will define what progress means for the autonomous company.

What you’ll do
  • Work with domain experts to identify the decisions, constraints, and outcomes that matter in a workflow.

  • Design tasks, evaluation environments, rubrics, and verifiers that test complete workflows, including ambiguous inputs and failure recovery.

  • Build data pipelines that preserve the relationship between context, decisions, actions, and outcomes, with clear provenance and appropriate access controls.

  • Create evaluation splits and review processes that limit leakage, expose distribution shifts, and catch regressions.

  • Study agent traces to identify recurring failures and turn the findings into better training data, tools, and reward signals.

  • Build repeatable evaluation reports that let researchers and engineers compare quality, reliability, latency, and cost.

What you’ll bring
  • Strong Python and practical experience building data or machine learning systems.

  • Evidence of designing useful evaluations, experiments, or datasets for models or complex software.

  • Sound statistical judgment and the ability to explain what a metric establishes and what it misses.

  • Care with data quality, reproducibility, and the handling of business information.

  • The ability to learn an unfamiliar domain and communicate clearly with both technical and operational colleagues.

Useful experience

Agent evaluation, reward modeling, human feedback systems, synthetic data, process mining, or research involving sequential decisions. Published work and open-source contributions are welcome; concrete examples of careful measurement matter just as much.

What success looks like

The team can tell whether a change improves real task outcomes, reproduce the result, and trace failures back to actionable causes.

Apply

Show us something you have built, investigated, or improved, and tell us what you learned. That might be a product, a research result, an open-source contribution, or work you can describe without disclosing confidential information. We welcome strong evidence of ability from a range of backgrounds.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Research — RL & Agents
Applied Research — RL & Agents

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 180,000
Applied Research — Forward Deployed
Applied Research — Forward Deployed

Human Intuition Inc. • New York (NY)

On-site
USD 150,000 - 190,000
Member of Technical Staff — Agent Systems
Member of Technical Staff — Agent Systems

Human Intuition Inc. • New York (NY)

On-site
USD 150,000 - 190,000
Member of Technical Staff — Full Stack
Member of Technical Staff — Full Stack

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Applied Researcher: Agent Evaluation & Data Pipelines
Applied Researcher: Agent Evaluation & Data Pipelines

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 180,000
Member of Technical Staff — Inference
Member of Technical Staff — Inference

Human Intuition Inc. • New York (NY)

On-site
USD 140,000 - 195,000
Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Member of Technical Staff, Agents
Member of Technical Staff, Agents

Engg • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Staff Product Manager (Evals) Palo Alto, California
Staff Product Manager (Evals) Palo Alto, California

Workato • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000