Judgment Labs — Full-Stack Engineer

davidjoseph-co

San Francisco (CA)

On-site

USD 200,000 - 300,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Full benefits
Equinox membership
Private chef

Job summary

Judgment Labs is seeking a Full-Stack Engineer to own the product experiences and agent infrastructure from data layer to UI, in a San Francisco, CA on-site role. You will work 5 days per week in-person, collaborating with the founding team to deliver scalable agent monitoring infrastructure.

The role requires 3–7 years of full-stack engineering, hands-on LLM/agent-building, and strong customer-facing communication, with relocation supported.

Qualifications

  • 3 to 7 years full-stack engineering experience.
  • Hands-on LLM or agent-building experience.
  • End-to-end production system ownership, data layer to UI.
  • Customer-facing communication ability (approximately 30% of role).
  • SF in-person 5 days per week, relocation supported.

Responsibilities

  • Own the product experiences and agent infrastructure spanning data layer to UI.
  • Shape large-scale parallel investigations across thousands of production traces.
  • Build hosted simulated environments for stateful evals and trajectory replay.
  • Design how engineers understand long traces, tool calls, decisions, and failures for quick insights.
  • Develop and maintain the SDK and terminal-first experience for agent-dev sessions.

Skills

Full-stack development
LLM/agent-building
Production systems ownership
Customer-facing communication

Tools

SDK development
LLM infrastructure
Agent tooling

Job description

Judgment Labs — Full-Stack Engineer

Type: Full-time | On-site | San Francisco, CA (5 days/week in-person)Compensation: $200,000–$300,000 base + equityHiring count: 2Visa sponsorship: Case-by-case for truly exceptional candidates (H-1B, O-1, OPT); primary scope is candidates who don't require sponsorshipReports to: Founding team (no named contact on role page)

About Judgment Labs

Judgment Labs builds infrastructure for Agent Behavior Monitoring (ABM). Where traditional observability logs exceptions and latency, ABM surfaces behavioral anomalies — instruction drift, context-retrieval loss — in scaled production environments. Hundreds of teams building autonomous agents rely on Judgment to understand how their systems behave post-deployment.

The team has raised $30M+ across two rounds in the past five months, backed by Lightspeed, SV Angel, Valor Equity Partners, Nova Global, Chris Manning, Michael Ovitz, Michael Abbott, Cory Levy, and Kevin Hartz. Under 20 people, shipping at 50+ company velocity, with Olympiad medalists, debate champions, and competitive athletes. Everyone is either an ex-founder or a founder-to-be.

Founded: N/A | Team size: <20 | Total funding: $30M+Industry: AI agent infrastructure / observabilityWebsite: judgmentlabs.aiOffice: San Francisco, CA

Why Candidates Should Join
  • Category-defining space: Building the ABM layer beyond traditional observability — how autonomous agents are understood in production.
  • Well-funded, moving fast: $30M+ across two rounds in five months, top-tier backers, under-20 team shipping at 50+ company pace.
  • Real product ownership: Not spec implementation — you talk to customers, define what to build, build it, and iterate end-to-end.
  • Elite, intense team: Olympiad medalists, debate champions, ex-founders; everyone builds like a founder.
  • Comp + perks: Up to $300K (up to $400K for exceptional AI-savvy FDE leads), full benefits, Equinox membership, private chef.
Intake Call Summary
  • No intake call transcript was provided on the role page. An intake video is available on the Contrario listing but was not captured here.
  • Key hiring signals inferred from the page: strong bias toward genuine software-engineering depth (not solutions/FDE-only backgrounds), agent/LLM fluency, and candidates for whom this product-engineering role is a genuine first choice rather than a fallback.
The Role

Own the product experiences and agent infrastructure that make the agent improvement loop legible and actionable for engineering teams. Spans the data layer to the UI — agent swarm interfaces, verification platforms, and the SDK layer that lets developers summon Judgment mid-development. ~30% customer-facing.

What You'll Be Doing
  • Shape how the Judgment Agent runs large-scale parallel investigations across thousands of production traces, merging failure modes, tool errors, regressions, and drift signals into a single actionable answer.
  • Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior.
  • Design how engineers understand long traces, tool calls, decisions, and failures — making a thousand-step reasoning trajectory legible in minutes.
  • Build the swarm UX so engineers can watch parallel investigations, redirect investigators going down the wrong path, and consume findings without reading hundreds of reports.
  • Own the improvement loop: workflows that turn production trajectories into datasets, judges, and regression checks so found-problem to verified-fix feels like one motion.
  • Build and maintain the SDK and terminal-first experience so agent-dev sessions can summon Judgment as a subagent mid-development.
  • Own platform infrastructure: workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.

Tech stack: Not specified on role page. Full-stack (data layer through UI), SDK development, terminal-first tooling, agent/LLM infrastructure.

Requirements
  • 3 to 7 years full-stack engineering experience
  • Hands-on LLM or agent-building experience
  • End-to-end production system ownership, data layer to UI
  • Customer-facing communication ability (approximately 30% of role)
  • SF in-person 5 days per week, relocation supported
Green Flags
  • Technical background (e.g. coding competitions, research experience) in addition to solutions experience
  • Genuine first choice for a product-engineering role with agent depth
  • Prior evals, observability, or behavior-monitoring product background
  • 4 to 5 years of experience with strong communication track record
  • Strong engineering pedigree from a solid product company
  • Founder background or clear founder-to-be signal
Red Flags
  • FDE is a fallback, not their top choice ("if I can't get X, I'll do FDE")
  • Low commitment signals: slow to book or drops off before the interview
  • Over-indexed on agent experience at the expense of engineering quality
  • Pure solutions or forward-deployed background without real software engineering depth
  • Big-tech candidate using this role as a backup
Role Details
  • Salary — $200,000–$300,000 (junior ~$200K at 3 yrs; mid-to-senior up to $300K at 3–7 yrs; exceptional AI-savvy FDE leads up to $400K case-by-case)
  • Equity — Yes (amount not specified)
  • On-site policy — SF in-person, 5 days/week; relocation supported
  • Visa sponsorship — Case-by-case for exceptional candidates (H-1B, O-1, OPT); main scope is candidates not requiring sponsorship
  • Employment type — Full-time
  • Location — San Francisco, CA

Benefits & perks: Full benefits package Equinox membership Private chef Competitive compensation Direct customer interface and influence on product roadmap

Screening Questions

None provided on the role page.

Interview Process

Stage 1 — Recruiter screen — Initial screen.Stage 2 — Evals FDE round — Evals / FDE-focused round.Stage 3 — Coding IQ round — Coding / problem-solving round.Stage 4 — FDE onsite — Onsite.Stage 5 — Offer ExtendedStage 6 — Candidate Hired — Candidate accepts and starts.

(Platform scheduling states — "Pending Approval," "booked evals round," "booked iq round" — omitted as pipeline artifacts.)

Ideal Companies & Backgrounds

Updated from role pageTarget product / infra / AI companies — Palantir, Databricks, Datadog, Cognition AI, Decagon, Sierra, Linear, Cursor, Ramp, Figma, Vercel, CockroachDB, Modal Labs, Anyscale, Runway, Applied Intuition, Anduril, Notion, Nomic AI, MotherDuck

Body callouts: solid product companies (e.g. Roblox, Snapchat, Tesla); top-tier infra (e.g. Databricks) is a bonus.

Note: several entries on the page (".Store," "Data Center ISH," "Retoolers," "Mercury Insurance," "SearchHounds," "Ponderosa Agency") appear to be logo-resolver artifacts rather than genuine target companies and are excluded above.

Ideal Candidate Profiles

For reference only — DO NOT CONTACT. No LinkedIn URLs were exposed on the role page. Yuval Danino Bhagyashri Badgujar Smriti Sridhar Elie Harik Kabeer Thockchom Sukrit Rao Sujan Rachuri Ishan Mehta Krrish Chawla Joseph Tey Aditya Tadimeti Sathvik Nallamalli Aliyan Ishfaq

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Judgment Labs — Agent Product Engineer
Judgment Labs — Agent Product Engineer

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 350,000
Full benefits package
Equity
Private chef
Init Intelligence — Senior/Staff Applied AI Engineer, Agent Harness
Init Intelligence — Senior/Staff Applied AI Engineer, Agent Harness

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 300,000
Meals in office
Health insurance
Unlimited PTO
Forward Deploy AI Engineer — Judgment Labs
Forward Deploy AI Engineer — Judgment Labs

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 300,000
Full benefits
Equinox membership
Private chef
Init Intelligence — Senior/Staff Founding Full-Stack Engineer, Agent Systems
Init Intelligence — Senior/Staff Founding Full-Stack Engineer, Agent Systems

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 300,000
Meals in office
Health insurance
Unlimited PTO
Growth Associate — AfterQuery
Growth Associate — AfterQuery

davidjoseph-co • San Francisco (CA)

On-site
USD 100,000 - 130,000
Specific Labs — Founding Strategic Projects Lead
Specific Labs — Founding Strategic Projects Lead

davidjoseph-co • San Francisco (CA)

On-site
USD 120,000 - 200,000
DoorDash meal stipend
Gym in building
Health insurance
Varick Agents — Staff / Senior Engineer
Varick Agents — Staff / Senior Engineer

davidjoseph-co • San Francisco (CA)

On-site
USD 225,000 - 300,000
Bluejay — Member of Technical Staff
Bluejay — Member of Technical Staff

davidjoseph-co • San Francisco (CA)

On-site
USD 160,000 - 220,000
Competitive equity
Varick Agents — Head of Engineering
Varick Agents — Head of Engineering

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 300,000
Jampack AI — Founding Engineer, Fullstack
Jampack AI — Founding Engineer, Fullstack

davidjoseph-co • New York (NY)

Hybrid
USD 160,000 - 220,000
Health insurance
Equity compensation