Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Sycamore is building the trusted agent operating system for the enterprise. Our small, engineering-led team works directly with Fortune 500 companies to deploy agentic apps with the security and control large organizations need.
You will help build the runtime, evaluation, and learning systems that power the platform, owning end-to-end slices from the agent runtime to environments and self-improvement, collaborating with product, infrastructure, and Forge teams.
Build the runtime, evaluation, and learning systems that every agent on the platform depends on.
Sycamore is building the trusted agent operating system for the enterprise. Our platform helps companies build, deploy, and orchestrate agentic apps that take on real operational work, with the security and control large organizations need.
We are a small, engineering-led team working directly with Fortune 500 enterprises. We have raised $65M from Coatue and Lightspeed, along with other investors and industry leaders.
Core AI owns the horizontal runtime, intelligence, and improvement capabilities that Product and Infrastructure both depend on, along with the Sycamore Forge experiences that make them usable. Engineers here own end-to-end slices, from cloud service and API design through the React interface. The team covers four related areas, and you do not have to pick one to apply. Tell us what you have built and we will work out the fit together.
Multi-turn sessions, model routing, tool execution, memory, durable workflows, and the APIs that expose them. The hard part is correctness across long horizons: surviving retries, provider interruptions, partial failure, and context growth without losing the thread.
An agent that writes and runs code needs somewhere to do it. That environment has to start fast, isolate genuinely untrusted execution, reach only the services it legitimately needs, and carry credentials it can use but never read. The same area owns the verification layer that decides whether what an agent produced actually works rather than only appears to.
Agent quality is genuinely hard to measure. A change that looks better on a handful of examples often is not, and a judge model can be confidently wrong in the same direction as the system it grades. This is where the offline suites, replay corpora, and statistical discipline that gate a release come from.
Every agent run produces evidence, and almost all of it is currently thrown away. This area turns it into improvement: structured trajectories at fleet scale, failure clusters surfaced from production rather than guessed at, and proposed changes that are versioned, measured, staged, and reversible.
Our current Core AI environment includes Python cloud services; React and TypeScript product surfaces in Sycamore Forge; asynchronous and streaming systems; typed APIs and data models; relational and vector data; durable workflows; protocol-based tool execution; multiple model providers; and cloud-native deployment.
We use coding agents, automated tests, traces, evaluations, browser automation, offline replay, cost and latency signals, and production feedback as part of everyday engineering. We are building toward a governed collect, learn, evaluate, and apply loop rather than a single monolithic training system.
This is context, not a checklist. We do not require previous experience with every language, framework, model provider, cloud platform, database, or infrastructure tool in our stack. Comparable experience building distributed runtimes, experimentation platforms, retrieval or recommendation systems, workflow engines, developer platforms, or production AI systems is highly relevant.
None of the following is required, and any of them is a strong signal for a particular area: production Kubernetes work deep enough to have written controllers rather than only configured them; a real understanding of isolation boundaries and what each one does and does not contain; comfort with statistics, including the ways an experiment can mislead you; trajectory or event data at scale, including the joins, lineage, and privacy handling that make it usable; or reinforcement learning, post-training, or large-scale experimentation. We care more about the systems you personally built, measured, and operated than a particular company, school, language, or model vendor.