product engineer, agent

Judgment Labs

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Judgment Labs looks for a Product Engineer who will own end-to-end production experiences for AI agents. You’ll shape how large-scale investigations run, and build verification environments with trajectory replay to ensure changes are correct and robust.

You’ll also design interfaces for engineers to understand agent behavior, while enabling a swarm-style view of parallel investigations and repeatable workflows.

Qualifications

  • Experience building and scaling end-to-end production systems.
  • Strong problem-solving skills in fast-changing, ambiguous environments.
  • Hands-on experience with LLMs or agents.
  • Comfort talking to customers to understand needs and translate into features.
  • Excellent communication across technical and non-technical audiences.

Responsibilities

  • Shape how Judgment Agent runs large-scale investigations across thousands of production traces and dimensions.
  • Build the platform for verifying agent changes with simulated environments and trajectory replay.
  • Design interfaces to show engineers what their agents did and why.
  • Create Swarm UX so multiple investigations can be followed and steered efficiently.
  • Develop workflows turning trajectories into datasets, judges, and regression checks.
  • Oversee the platform aspects: workspaces, roles, permissions, billing, usage.
  • Provide an SDK/terminal-first experience for agent sessions to summon Judgment during development.

Skills

End-to-end production systems
LLMs / agents
Customer-facing communication
Problem solving
Rapid iteration

Job description

Product Engineer — Agents Job Description
Product Engineer
The Role

Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:

  1. We ingest everything your agents do in production: traces, tool calls, decisions, outcomes

  2. Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals

  3. Teams close the loop, shipping agent improvements validated against real production evidence

You’ll build the product experiences that make this loop legible, and you’ll build the agents that run it. This is not a role where you implement specs handed down. You’ll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it’s great.

What You Will Accomplish
  • Judgment Agent: Shape how the Judgment Agent runs large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer.

  • Verification: Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior changes.

  • Agent investigation interfaces: Design how engineers understand what their agents did and why. Long traces, tool calls, decisions, failures. What does debugging look like when the "program" is a reasoning loop? How do you make a thousand-step trajectory legible in minutes?

  • Swarm UX: A hundred parallel investigations is useless if engineers can't follow them. Design how humans watch a swarm work, redirect investigators that go down the wrong path, and consume findings without reading a hundred reports.

  • The improvement loop: Build the workflows that turn production trajectories into datasets, judges, and regression checks, so the path from "found a problem" to "verified a fix" feels like one motion.

  • The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.

  • Judgment everywhere agents are built: An SDK and terminal-first experience so Claude Code, Codex, and OpenCode sessions can summon Judgment as a subagent mid-development.

What You’ll Bring
  • Experience building and scaling end-to-end production systems, from data layer to UI

  • Strong technical problem-solving skills, especially in fast-changing, ambiguous environments

  • A builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship

  • Hands-on experience building with LLMs or agents, or the drive to get there fast

  • Comfort working directly with customers to understand their needs and solve real-world problems

  • Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product engineer, full stack
Product engineer, full stack

Judgment Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Product engineer, full stack
Product engineer, full stack

Judgment Labs • San Francisco (CA)

On-site
USD 140,000 - 210,000
Product Engineer - AI Judgment Platform
Product Engineer - AI Judgment Platform

Judgment Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer
Research Engineer

Ersilia • San Francisco (CA)

On-site
USD 150,000 - 190,000
Full-Stack Engineer - AI Agent Platform & Debug UX
Full-Stack Engineer - AI Agent Platform & Debug UX

Judgment Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Judgment Labs — Agent Product Engineer
Judgment Labs — Agent Product Engineer

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 350,000
Full benefits package
Equity
Private chef
Agent Product Engineer — Build 0→1 AI Agents (SF On-site)
Agent Product Engineer — Build 0→1 AI Agents (SF On-site)

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 350,000
Full benefits package
Equity
Private chef
Product Manager, Agent Experience
Product Manager, Agent Experience

CrewAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Builder
AI Builder

Curb Mobility • New York (NY)

On-site
USD 100,000 - 140,000
Product Manager, AI Agents
Product Manager, AI Agents

Atom Matrix Innovation LLC • San Francisco (CA)

On-site
USD 150,000 - 210,000