Stand out for this role — generate a tailored resume and cover letter in about a minute.
Sycamore is building the trusted agent operating system for the enterprise, focusing on secure execution substrates for enterprise agents, customer applications, and Sycamore’s control plane. We seek an experienced software/infrastructure engineer to design and operate production-grade control-plane services, tenant isolation, and scalable database platforms.
You will own end-to-end reliability across deployments, assist with customer deployments, and contribute to a high-velocity, on-call
Build and operate the secure execution substrate for enterprise agents, customer applications, and Sycamore’s control plane.
Sycamore is building the trusted agent operating system for the enterprise. Our platform helps companies build, deploy, and orchestrate agentic apps that take on real operational work, with the security and control large organizations need.
We are a small, engineering-led team working directly with Fortune 500 enterprises. We have raised $65M from Coatue and Lightspeed, along with other investors and industry leaders.
This is software engineering for consequential distributed systems, not cloud administration: control-plane services and Kubernetes controllers, identity and network boundaries, tenant-isolated data, and failures that cross application, cluster, database, and customer-network layers. Infrastructure is the foundation beneath Product and Core AI, so an error here can affect every application and customer. The team covers three areas. They share an on-call rotation and a great many failures, so nobody works in only one.
The platform and control plane. Kubernetes-based serving, the sandboxes agents execute in, deployment and rollback, networking and identity, and the operators that reconcile all of it. This is where blast radius is decided.
The data platform. This is database platform engineering, not analytics: there is no warehouse, no dbt, and no pipeline waiting for an owner. What we have is a growing fleet of PostgreSQL databases carrying enterprise customer data under a genuine isolation requirement, hundreds of migrations across many trees and authors, a connection budget that is a real ceiling, retention obligations measured in years, and a hybrid full-text and vector retrieval path in the hot path of every agent conversation. The problems are fleet-shaped: a migration is easy, and applying it safely across every tenant database with per-tenant failure isolation and no downtime is not.
Reliability. How the platform behaves over time: service level objectives and error budgets that mean something, an observability estate managed as code, promotion gates that decide whether a release reaches production, capacity and cold-start performance, and incident response from page through postmortem to the structural fix. We have a large monitor fleet and an enforced metric taxonomy, but few service level objectives, no error budgets, no burn-rate alerting, and a promotion gate that warns rather than blocks. Building that practice is the work, not a side project within it.
Our current infrastructure environment includes GCP, Kubernetes and GKE, Knative, Envoy, Cloudflare, Pulumi, Helm, BuildKit, PostgreSQL and Cloud SQL, object storage, container registries, workload identity, Python and Go control-plane services, Kubernetes operators, GitHub Actions, and production observability.
Alert triage is substantially agent-driven, with automated investigation, ticket filing, and a decay process that nominates monitors nobody has acted on for deletion. We also support customer-controlled deployment requirements and maintain more than one application-serving path during platform migrations.
This is context, not a checklist. We do not require previous experience with every cloud product or tool. We care about deep systems fundamentals, the ability to learn unfamiliar infrastructure, and evidence that you have personally operated and recovered important multi-tenant systems.
We do not expect all of the following from one person, and any of it is a strong signal for a particular area: deep PostgreSQL expertise, including connection management, query planning, online schema change, and multi-tenant isolation as an enforced guarantee rather than an application convention; Kubernetes performance and capacity depth; or incident command experience and postmortems that produce structural change rather than an action item nobody does. We care more about what you personally built, operated, and recovered than a particular school or vendor certification.