Member of Technical Staff: Agent Runtime

ego AI (YC W24)

San Francisco (CA)

Hybrid

USD 150,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Ego AI, a YC-backed applied AI research lab in San Francisco, is hiring for a Member of Technical Staff to own the agent runtime. The role focuses on building a durable, extensible agent harness with a robust plan-tool-observe loop, per-user durable agents, and strong session management.

Hybrid work supported. You will design a per-user memory-enabled agent system, implement a tool registry, and ensure reliable execution from plan through tool calls, with emphasis on safety, observability, and

Qualifications

  • 3+ years building backend or infrastructure systems with production ownership
  • Hands-on experience building LLM agent systems: tool-calling loops, defensive handling of model output, and budgets
  • Deep grasp of durable execution: checkpointing, idempotency, outbox/step-log patterns
  • Deliberate context-management strategy for long-horizon tasks and trade-offs
  • Strong API and data-model design; persistence behind clean seams
  • Clear written communication and architecture documentation

Responsibilities

  • Own the agent harness and core decision loop (plan → tool call → observe → repeat) as a maintainable state machine
  • Build checkpointed, resumable sessions with durable logs and idempotent tooling
  • Develop compaction, running summaries, and retrieval over older state for long-horizon tasks
  • Design a tool registry with name, JSON schema, and handler; ensure safe schema validation
  • Ship deployable systems with local bring-up, persistence interfaces (Postgres/SQLite/Cloudflare storage), and per-user durable agents
  • Support engineers by reviewing designs and extending the harness

Skills

Backend systems
LLM agent systems
Durable execution
Context management
API & data-model design
Clear documentation

Tools

Cloudflare Durable Objects
Postgres/SQLite

Job description

Member of Technical Staff: Agent Runtime (Ego) Location: San Francisco, CA (hybrid)



  • Reports to: Founding team


About The Role

Ego AI is a YC-backed applied AI research lab building the behavioral infrastructure for AI companions and agents. We work at the intersection of real-time conversational instincts, memory, and persistent identity; the layer that makes AI feel genuinely alive. We're a small, fast-moving team and we're defining a new category of human-AI interaction.


What you'll do

Own the agent harness. Design and maintain our core agentic loop (plan → tool call → observe → repeat → finalize) as a small, legible state machine. Errors are first-class: malformed tool calls, hallucinated tool names, and throwing tools get fed back to the model. They never crash the loop.


Make execution durable. Build checkpointed, resumable sessions backed by a database: step logs with pending/committed status, idempotency keys on side-effecting tools, and a clear account of the at-least-once vs exactly-once boundary. A crash between the LLM call and the tool call, or mid-compaction, must never corrupt a session or double-fire a POST.


Solve long-horizon context. Own our compaction strategy: pinned goals, running summaries, sliding windows of recent turns, and retrieval over older state. A 30+-step task should finish aimed at the original goal, within budget.


Design for extensibility. Ship a tool registry (name + JSON schema + handler) that teammates and partners can extend without touching the loop. Validate schemas before handlers run. Keep model adapters swappable in one place.


Ship deployable systems. Deliver one-command local bring-up, persistence behind an interface (Postgres SQLite Cloudflare storage), clean HTTP APIs, and secrets hygiene. Build per-user durable agent instances that wake on triggers (cron, webhooks, inbound events). Cloudflare Durable Objects experience is a real plus.


Support the team. Unblock product engineers building on the harness, pair on extensions, and review agent-adjacent designs.


Projects this hire will own


  • Productionize the agent runtime ("Ronin core"). Take our harness from working prototype to the shared runtime every Ego product sits on: durable step log, budget enforcement, compaction, tool registry, and observability/tracing.

  • Per-user durable personal agents. Each user gets a long-lived agent instance with its own memory, triggers, and budget. It wakes on webhooks or cron, survives restarts, and stays isolated per tenant. Likely on Cloudflare Durable Objects or an equivalent single-writer model.

  • Evaluation and reliability harness. Build repeatable crash tests (kill -9 mid-task and mid-compaction), race tests on concurrent session writes, cheap long-horizon stubs that exercise compaction without burning tokens, and per-step tracing.


What we're looking for

Must have


  • 3+ years building backend or infrastructure systems, with real production ownership of stateful services or relevant projects that they have worked on.

  • Hands-on experience building LLM agent systems: tool-calling loops, defensive handling of model output, and step/token/cost budgets.

  • Deep grasp of durable execution: checkpointing, idempotency, outbox/step-log patterns, and the ability to reason precisely about failure windows (what if it dies after the side effect but before it is recorded?).

  • A deliberate context-management strategy for long-horizon tasks, and the ability to defend its trade-offs rather than only describe it.

  • Strong API and data-model design; persistence behind clean seams; comfort being judged on docker compose up working cold from a README.

  • Clear written communication. Architecture docs a reviewer understands before reading the code. Honest scoping (what you cut and why).


Nice to have


  • Cloudflare Workers / Durable Objects / D1 in production.

  • Multi-model experience (DeepSeek, Claude, open-weight models) and eval tooling.

  • Long-term memory systems for agents (personalization layers, retrieval over older state, hierarchical summaries).

  • Concurrency safety on shared session state; streaming output; structured per-step observability.


Use AI tools freely. We do. You will walk us through the design and defend every trade-off, so own each decision in the repo.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agent Builder
Agent Builder

Aisle • New York (NY)

On-site
USD 140,000 - 180,000
Senior AI Engineer - Forward Deployed (FDE)
Senior AI Engineer - Forward Deployed (FDE)

Moring Ai • Atlanta (GA)

On-site
USD 150,000 - 210,000
AI Engineer, Agent Builder (Remote).
AI Engineer, Agent Builder (Remote).

Catalyst Wayfare • United States

Remote
USD 120,000 - 170,000
Agent Harness Engineer
Agent Harness Engineer

Axiom • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
AI Engineer
AI Engineer

Aspire Software • United States

On-site
USD 120,000 - 180,000
Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Data & Machine Learning Engineer
Data & Machine Learning Engineer

Filmore • Austin (TX)

On-site
USD 140,000 - 180,000
Evals Lead
Evals Lead

Aslan • Washington

On-site
USD 120,000 - 180,000
Agent Builder
Agent Builder

BAM Ventures • New York (NY)

On-site
USD 140,000 - 190,000
AI Engineer
AI Engineer

Valsoft Corporation • Northern (KY)

Hybrid
USD 120,000 - 180,000