Get more replies from employers
Send a job-specific resume in minutes.
Judgment Labs is seeking a Full-Stack Engineer to own the product experiences and agent infrastructure from data layer to UI, in a San Francisco, CA on-site role. You will work 5 days per week in-person, collaborating with the founding team to deliver scalable agent monitoring infrastructure.
The role requires 3–7 years of full-stack engineering, hands-on LLM/agent-building, and strong customer-facing communication, with relocation supported.
Type: Full-time | On-site | San Francisco, CA (5 days/week in-person)Compensation: $200,000–$300,000 base + equityHiring count: 2Visa sponsorship: Case-by-case for truly exceptional candidates (H-1B, O-1, OPT); primary scope is candidates who don't require sponsorshipReports to: Founding team (no named contact on role page)
Judgment Labs builds infrastructure for Agent Behavior Monitoring (ABM). Where traditional observability logs exceptions and latency, ABM surfaces behavioral anomalies — instruction drift, context-retrieval loss — in scaled production environments. Hundreds of teams building autonomous agents rely on Judgment to understand how their systems behave post-deployment.
The team has raised $30M+ across two rounds in the past five months, backed by Lightspeed, SV Angel, Valor Equity Partners, Nova Global, Chris Manning, Michael Ovitz, Michael Abbott, Cory Levy, and Kevin Hartz. Under 20 people, shipping at 50+ company velocity, with Olympiad medalists, debate champions, and competitive athletes. Everyone is either an ex-founder or a founder-to-be.
Founded: N/A | Team size: <20 | Total funding: $30M+Industry: AI agent infrastructure / observabilityWebsite: judgmentlabs.aiOffice: San Francisco, CA
Own the product experiences and agent infrastructure that make the agent improvement loop legible and actionable for engineering teams. Spans the data layer to the UI — agent swarm interfaces, verification platforms, and the SDK layer that lets developers summon Judgment mid-development. ~30% customer-facing.
Tech stack: Not specified on role page. Full-stack (data layer through UI), SDK development, terminal-first tooling, agent/LLM infrastructure.
Benefits & perks: Full benefits package Equinox membership Private chef Competitive compensation Direct customer interface and influence on product roadmap
None provided on the role page.
Stage 1 — Recruiter screen — Initial screen.Stage 2 — Evals FDE round — Evals / FDE-focused round.Stage 3 — Coding IQ round — Coding / problem-solving round.Stage 4 — FDE onsite — Onsite.Stage 5 — Offer ExtendedStage 6 — Candidate Hired — Candidate accepts and starts.
(Platform scheduling states — "Pending Approval," "booked evals round," "booked iq round" — omitted as pipeline artifacts.)
Updated from role pageTarget product / infra / AI companies — Palantir, Databricks, Datadog, Cognition AI, Decagon, Sierra, Linear, Cursor, Ramp, Figma, Vercel, CockroachDB, Modal Labs, Anyscale, Runway, Applied Intuition, Anduril, Notion, Nomic AI, MotherDuck
Body callouts: solid product companies (e.g. Roblox, Snapchat, Tesla); top-tier infra (e.g. Databricks) is a bonus.
Note: several entries on the page (".Store," "Data Center ISH," "Retoolers," "Mercury Insurance," "SearchHounds," "Ponderosa Agency") appear to be logo-resolver artifacts rather than genuine target companies and are excluded above.
For reference only — DO NOT CONTACT. No LinkedIn URLs were exposed on the role page. Yuval Danino Bhagyashri Badgujar Smriti Sridhar Elie Harik Kabeer Thockchom Sukrit Rao Sujan Rachuri Ishan Mehta Krrish Chawla Joseph Tey Aditya Tadimeti Sathvik Nallamalli Aliyan Ishfaq