San Francisco, CA - On-site (5 days/week) - Full-time
Compensation: $200,000-$300,000 + equity
About the Company
Our client is a well-funded San Francisco startup (Series A, $30M+ raised) building infrastructure that monitors and evaluates how AI agents behave in production. Their platform goes beyond traditional observability - surfacing behavioral issues like instruction drift and context-retrieval loss so teams can understand and continuously improve their autonomous agents after deployment. Hundreds of teams already rely on them. The team is small, elite, and moving extremely fast.
Founded: recent-stage startup - Under 20 people - Industry: AI agent infrastructure / observability
The Role
This is a genuine full-ownership product-engineering role, not spec implementation. You'll talk to customers, decide what to build, build it, and iterate until it's great - across the entire stack, from the data layer to the UI. Expect to work on agent-investigation interfaces, verification and evaluation platforms, and a developer SDK. About 30% of the role is customer-facing.
What you'll be doing
- Build interfaces that make large-scale, parallel agent investigations legible - turning thousands of production traces and a wide range of failure signals into a single actionable answer.
- Build the platform for verifying agent changes: simulated environments for stateful evals, trajectory replay, and monitors for unintended behavior.
- Make long, multi-step reasoning trajectories understandable in minutes.
- Own the improvement loop end to end, turning production data into datasets, judges, and regression checks.
- Build and maintain the SDK and a terminal-first developer experience.
- Own core platform infrastructure: workspaces, permissions, billing, usage, and limits.
Tech stack: Full-stack (data layer through UI), SDK development, terminal-first tooling, agent/LLM infrastructure.
Requirements
- 3-7 years of full-stack engineering experience, comfortable owning systems from the data layer through the UI
- Hands-on experience building with LLMs or AI agents (or a clear, demonstrated drive to ramp fast)
- A track record of shipping and scaling production systems end to end
- Comfortable working directly with customers to understand needs and solve real problems (~30% of the role)
- Strong problem-solving in fast-moving, ambiguous environments
- Clear, direct, persuasive communication across technical and non-technical audiences
- Based in San Francisco or open to relocating; in-office 5 days a week
Nice to Haves
- Prior work on evals, observability, or agent behavior-monitoring products
- Strong engineering pedigree from a respected product company; top-tier infrastructure experience is a bonus
- Prior forward-deployed or solutions engineering experience
- Early-stage / 0-to-1 startup experience as an early hire
- A technical edge beyond delivery - coding competitions, research, or similar
- Founder background or a clear founder-to-be trajectory
- Full-stack comfort with real backend and infrastructure depth
Why Join
- Own a real product surface end to end - from customer conversations to shipped features.
- Join a small, high-velocity, well-backed team working at the frontier of AI agent reliability.
- Competitive compensation up to $300K + equity, with a strong benefits and perks package.
Details
- Location - San Francisco, CA
- Work policy - On-site, 5 days/week; relocation supported
- Compensation - $200,000-$300,000 + equity
- Visa sponsorship - Considered case-by-case for exceptional candidates; primary scope is candidates who don't require sponsorship
- Employment type - Full-time