Member of Technical Staff, Infrastructure

Sycamore

Palo Alto (CA)

On-site

USD 180,000 - 260,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Sycamore is building the trusted agent operating system for the enterprise, focusing on secure execution substrates for enterprise agents, customer applications, and Sycamore’s control plane. We seek an experienced software/infrastructure engineer to design and operate production-grade control-plane services, tenant isolation, and scalable database platforms.

You will own end-to-end reliability across deployments, assist with customer deployments, and contribute to a high-velocity, on-call

Qualifications

  • 5-12 years of software or infrastructure engineering experience.
  • Experience building production control-plane software, operators, platform services, or developer infrastructure, not only configuring vendor products.
  • Experience with infrastructure-as-code and delivery systems, including state management, drift detection, reproducible builds, staged rollout, and rollback.
  • Evidence of owning consequential production systems through incidents, capacity changes, recovery exercises, and reliability improvements.
  • A security-minded approach to tenant isolation, credentials, access control, software supply chains, and customer data.
  • The ability to write maintainable production software and automated tests for infrastructure behavior.
  • AI-native mindset: using coding agents and models as force multipliers while owning critical operational decisions.
  • Clear communication and high EQ when interacting with CTOs, security teams, and engineers.
  • Comfort with startup ambiguity and broad ownership, with occasional customer travel.

Responsibilities

  • Build and operate Kubernetes-based serving and sandbox platforms for agents, internal services, and customer applications.
  • Develop infrastructure-as-code, Helm charts, provisioning services, admission controls, and operators with safe change controls and rollout procedures.
  • Own the database fleet: schema topology, migrations, provisioning, rotation, and isolation guarantees.
  • Define service level indicators and objectives; manage observability estate as code: monitors, dashboards, tracing, and metric taxonomy.
  • Lead incident response from detection to postmortem and structural fixes.
  • Support hosted and customer-controlled deployment patterns with consistent operational model.
  • Travel to customer sites when collaboration improves deployment or incident outcomes.

Skills

5-12 years
Control-plane software
IaC & delivery
Production ownership
Tenant isolation
Automated tests
AI-native approach
Strong communication
Startup ambiguity

Job description

Build and operate the secure execution substrate for enterprise agents, customer applications, and Sycamore’s control plane.

About Sycamore

Sycamore is building the trusted agent operating system for the enterprise. Our platform helps companies build, deploy, and orchestrate agentic apps that take on real operational work, with the security and control large organizations need.

We are a small, engineering-led team working directly with Fortune 500 enterprises. We have raised $65M from Coatue and Lightspeed, along with other investors and industry leaders.

Where you could focus

This is software engineering for consequential distributed systems, not cloud administration: control-plane services and Kubernetes controllers, identity and network boundaries, tenant-isolated data, and failures that cross application, cluster, database, and customer-network layers. Infrastructure is the foundation beneath Product and Core AI, so an error here can affect every application and customer. The team covers three areas. They share an on-call rotation and a great many failures, so nobody works in only one.

The platform and control plane. Kubernetes-based serving, the sandboxes agents execute in, deployment and rollback, networking and identity, and the operators that reconcile all of it. This is where blast radius is decided.

The data platform. This is database platform engineering, not analytics: there is no warehouse, no dbt, and no pipeline waiting for an owner. What we have is a growing fleet of PostgreSQL databases carrying enterprise customer data under a genuine isolation requirement, hundreds of migrations across many trees and authors, a connection budget that is a real ceiling, retention obligations measured in years, and a hybrid full-text and vector retrieval path in the hot path of every agent conversation. The problems are fleet-shaped: a migration is easy, and applying it safely across every tenant database with per-tenant failure isolation and no downtime is not.

Reliability. How the platform behaves over time: service level objectives and error budgets that mean something, an observability estate managed as code, promotion gates that decide whether a release reaches production, capacity and cold-start performance, and incident response from page through postmortem to the structural fix. We have a large monitor fleet and an enforced metric taxonomy, but few service level objectives, no error budgets, no burn-rate alerting, and a promotion gate that warns rather than blocks. Building that practice is the work, not a side project within it.

What you will do
  • Build and operate Kubernetes-based serving and sandbox platforms for agents, internal services, and customer applications.
  • Develop infrastructure-as-code, Helm charts, provisioning services, admission controls, and operators that reconcile environments safely, along with the change controls, drift detection, release gates, and fail-safe rollout procedures around them.
  • Own the database fleet: schema topology and migration safety, provisioning and credential rotation, pooling and connection economics, isolation guarantees, retention and partitioning, and the search and embedding path.
  • Define service level indicators and objectives, establish error budgets, and own the observability estate as code: monitors, dashboards, tracing, metric taxonomy, burn-rate alerting, and the policy that keeps it from decaying.
  • Lead incident response from detection through containment, root cause, postmortem, and the structural fix.
  • Support hosted and customer-controlled deployment patterns while keeping the operational model consistent and auditable.
  • Travel to customer sites when working alongside a customer’s engineering or security team will materially improve a deployment or incident outcome.
The environment you will work in

Our current infrastructure environment includes GCP, Kubernetes and GKE, Knative, Envoy, Cloudflare, Pulumi, Helm, BuildKit, PostgreSQL and Cloud SQL, object storage, container registries, workload identity, Python and Go control-plane services, Kubernetes operators, GitHub Actions, and production observability.

Alert triage is substantially agent-driven, with automated investigation, ticket filing, and a decay process that nominates monitors nobody has acted on for deletion. We also support customer-controlled deployment requirements and maintain more than one application-serving path during platform migrations.

This is context, not a checklist. We do not require previous experience with every cloud product or tool. We care about deep systems fundamentals, the ability to learn unfamiliar infrastructure, and evidence that you have personally operated and recovered important multi-tenant systems.

What we are looking for
  • 5-12 years of software or infrastructure engineering experience. We will make exceptions for exceptional people in either direction.
  • Experience building production control-plane software, operators, platform services, or developer infrastructure, not only configuring vendor products.
  • Experience with infrastructure-as-code and delivery systems, including state management, drift detection, reproducible builds, staged rollout, and rollback.
  • Evidence of owning consequential production systems through incidents, capacity changes, recovery exercises, and reliability improvements.
  • A security-minded approach to tenant isolation, credentials, access control, software supply chains, and customer data.
  • The ability to write maintainable production software and automated tests for infrastructure behavior.
  • AI-native. You use coding agents and modern models as a force multiplier while still understanding and owning every critical operational decision.
  • Clear communication and high EQ. You can work with a customer’s CTO, security team, and infrastructure engineers without losing technical depth.
  • Comfort with startup ambiguity, broad ownership, and occasional customer travel.

We do not expect all of the following from one person, and any of it is a strong signal for a particular area: deep PostgreSQL expertise, including connection management, query planning, online schema change, and multi-tenant isolation as an enforced guarantee rather than an application convention; Kubernetes performance and capacity depth; or incident command experience and postmortems that produce structural change rather than an action item nobody does. We care more about what you personally built, operated, and recovered than a particular school or vendor certification.

Interview process
  • A 30-minute introductory conversation.
  • Two 60-minute technical interviews, one focused on systems design and one on coding.
  • A take-home assignment where you build and present a real solution using the tools you would use on the job.
Why join
  • Build the secure execution foundation for agents doing real work in large enterprises.
  • Solve systems problems spanning Kubernetes, control planes, identity, networking, data, delivery, and recovery.
  • Turn difficult customer deployment constraints into a platform that scales across environments.
  • Work on infrastructure where software design and operational judgment matter equally.
  • Join early enough to define how the engineering team operates and grows.
  • Receive competitive cash compensation and meaningful equity in the company you are helping build.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Infrastructure Intern Palo Alto
Member of Technical Staff, Infrastructure Intern Palo Alto

Sycamore • Palo Alto (CA), Northern (KY)

Hybrid
USD 55,000 - 96,000
Member of Technical Staff, Infrastructure Intern
Member of Technical Staff, Infrastructure Intern

Sycamore • Palo Alto (CA)

On-site
USD 42,000 - 66,000
Domain Expert, Semiconductors & Infrastructure
Domain Expert, Semiconductors & Infrastructure

Sycamore • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Member of Technical Staff, Products
Member of Technical Staff, Products

Sycamore • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Domain Expert, Semiconductors & Infrastructure Palo Alto
Domain Expert, Semiconductors & Infrastructure Palo Alto

Sycamore • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Equity
Member of Technical Staff, Core AI
Member of Technical Staff, Core AI

Sycamore • Palo Alto (CA)

On-site
USD 170,000 - 260,000
Equity
Competitive compensation
Deployment Strategist
Deployment Strategist

Sycamore • Palo Alto (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Member of Technical Staff, AI Services
Member of Technical Staff, AI Services

Sycamore • Palo Alto (CA)

On-site
USD 180,000 - 300,000
Member of Technical Staff, Applied AI
Member of Technical Staff, Applied AI

Sycamore • Palo Alto (CA), Northern (KY)

On-site
USD 200,000 - 260,000
Member of Technical Staff, Intern
Member of Technical Staff, Intern

Sycamore • Palo Alto (CA)

On-site
USD 54,000 - 67,000