Senior/Staff AI Infrastructure Engineer

Echelon

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
401(k)
Team meals
Unlimited PTO

Job summary

Echelon in San Francisco is building the AI platform for BizOps. We seek an engineer to design and operate sandbox lifecycle systems for agent code execution, browser work, document processing, and tool use. You will optimize startup time, ensure tenant isolation, durability of long-running jobs, and scale orchestration as workloads grow.

Strong Linux, Kubernetes, IaC, and cloud experience required. Join an early-stage company shipping production systems, working directly with founders and

Qualifications

  • 5+ years building production infrastructure, distributed systems, or execution runtimes.
  • Hands-on ownership of containerized workloads in a multi-tenant production environment.
  • Strong Linux systems knowledge across processes, filesystems, networking, and performance.

Responsibilities

  • Design and operate sandbox lifecycle systems for agent code execution, browser work, document processing, and tool use.
  • Improve sandbox startup time, density, scheduling, warm pools, caching, and resource utilization.
  • Build fast, reliable filesystem primitives for episodic and persistent agent state and large artifacts.
  • Enforce tenant isolation, network policy, secrets boundaries, quotas, and least-privilege access.
  • Make long-running agent jobs durable through checkpointing, retries, idempotency, cancellation, and recovery.
  • Scale orchestration and control-plane services with growth and bursts.
  • Build observability for resource pressure, execution failures, queue health, and end-to-end latency.

Skills

Production infra
Distributed systems
Container workloads
Kubernetes
Go
Rust
Performance tuning

Tools

Terraform
AWS

Job description

About Echelon

Echelon is building the AI platform for Business Operations. Our goal is to automate the knowledge work of BizOps so one exceptional operator can deliver the leverage of an entire team, as Ramp and Rippling have done for Finance and HR.

Our platform combines passive process mining with a living ontology that maps how work happens across people, systems, documents, decisions, and outcomes. It connects structured and unstructured data trapped in fragmented, duplicative, and legacy enterprise systems, then turns that context into automated workflows and AI agents.

We are an early-stage company tackling a difficult technical problem at high speed. Engineers work directly with founders and customers, make decisions with incomplete information, ship production systems, and own the results. The pace, rate of change, and standards are high.

The mandate

Build the secure execution substrate for Echelon's agents. You will scale ephemeral sandboxes, optimize filesystem and startup performance, and make long-running agent work durable under load without weakening tenant isolation or operational control.

What you'll own
  • Design and operate sandbox lifecycle systems for agent code execution, browser work, document processing, and tool use.
  • Improve sandbox startup time, density, scheduling, warm pools, caching, and resource utilization.
  • Build fast, reliable filesystem primitives for episodic and persistent agent state, large artifacts, and concurrent workloads.
  • Enforce tenant isolation, network policy, secrets boundaries, quotas, and least-privilege access.
  • Make long-running agent jobs durable through checkpointing, retries, idempotency, cancellation, and recovery.
  • Scale orchestration and control‑plane services through rapid workload growth and unpredictable bursts.
  • Build observability for resource pressure, execution failures, queue health, noisy neighbors, cost, and end‑to‑end latency.
  • Run load tests, capacity plans, failure drills, and incident reviews; fix root causes rather than adding fragile workarounds.
What you bring
  • 5+ years building production infrastructure, distributed systems, developer platforms, or execution runtimes.
  • Hands‑on ownership of containerized or virtualized workloads in a multi‑tenant production environment.
  • Strong Linux systems knowledge across processes, filesystems, networking, resource isolation, and performance debugging.
  • Experience with Kubernetes or a comparable scheduler, infrastructure as code, and cloud primitives on AWS or Azure.
  • Strong programming ability in Go, Rust, TypeScript, Python, or another systems‑oriented language.
  • Experience designing for retries, idempotency, backpressure, load shedding, observability, and safe rollouts.
  • Security instincts appropriate for executing untrusted or model‑generated work.
Useful experience
  • Firecracker, gVisor, Kata Containers, namespaces/cgroups, seccomp, or eBPF.
  • Sandbox products such as E2B, Modal, Fly Machines, or custom episodic compute platforms.
  • FUSE, overlay filesystems, content‑addressed storage, snapshotting, distributed caches, or object storage.
  • Agent runtimes, code interpreters, browser automation, remote development environments, or CI execution systems.
  • BYOC, private networking, customer‑managed deployments, or enterprise security reviews.
  • Inngest, Temporal, Kafka, NATS, or other durable workflow and event systems.
Benefits
  • Fully covered health, dental, and vision insurance.
  • 401(k) plan.
  • Team lunches and dinners in the office.
  • Unlimited PTO.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Staff AI Infrastructure Engineer
Senior/Staff AI Infrastructure Engineer

Socket.dev • San Francisco (CA)

On-site
USD 120,000 - 190,000
Health insurance
401(k) plan
Team meals in office
+1
Senior/Staff Applied AI Engineer
Senior/Staff Applied AI Engineer

Socket.dev • San Francisco (CA)

On-site
USD 190,000 - 260,000
Health, dental, and vision insurance
401(k) plan
Team lunches and dinners in the office
+1
Senior/Staff Product Engineer
Senior/Staff Product Engineer

Echelon • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
401(k)
Team meals
+1
Senior AI Infra Engineer: Scalable Sandboxes & Agents
Senior AI Infra Engineer: Scalable Sandboxes & Agents

Echelon • San Francisco (CA)

On-site
USD 180,000 - 280,000
Health insurance
Dental insurance
Vision insurance
+3
Backend Engineer (Systems)
Backend Engineer (Systems)

GenseeAI Inc. • California (MO)

On-site
USD 100,000 - 130,000
Competitive compensation in cash + equity
Ownership over projects
Direct influence on architecture and product direction
Platform Engineer - AI Infrastructure
Platform Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 180,000 - 250,000
Significant freedom and ownership in project development
Work on challenging problems related to ultra-low latency
Join a high-growth environment
Software Engineer, Infrastructure & Reliability
Software Engineer, Infrastructure & Reliability

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Engineering Team Lead, Platform
Engineering Team Lead, Platform

E2B • San Francisco (CA)

Hybrid
USD 190,000 - 240,000
Healthcare
Vision & Dental insurance
Unlimited PTO
+2
Backend Engineer
Backend Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 100,000 - 140,000
Health coverage
Opportunity to work with leading AI research labs
AI Infrastructure Engineer, Sandbox Platform
AI Infrastructure Engineer, Sandbox Platform

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 180,000 - 225,000