Founding Engineer (Physical AI Infrastructure)

a16z-speedrun

San Francisco (CA)

On-site

USD 190,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

a16z-speedrun is seeking an experienced systems engineer to build and operate the control and data planes for physical AI across public clouds, customer-owned infrastructure, robotics labs, and edge fleets. You will design orchestration across CPUs/GPUs, Kubernetes, bare metal, on-prem, and edge devices, while ensuring reliability for long-running workloads and complex data pipelines.

Role emphasizes production-grade Rust/Go/Python software, strong debugging across app, container, scheduler,

Qualifications

  • Built or operated large-scale distributed systems with reliability.
  • Strong experience with Kubernetes, IaC, networking, storage, security, and cloud architecture.
  • Understanding of GPU infrastructure, distributed training, batch systems, MLOps, or HPC.
  • Worked with high-volume streaming, time-series, image, video, or scientific data.
  • Production software in Rust, Go, Python, or similar systems language.
  • Able to debug across app, container, scheduler, network, host, GPU, and storage.
  • Care about internal architecture and developer experience for customers.

Responsibilities

  • Build a unified execution platform for simulation, policy evaluation, model training, data processing, and deployment.
  • Design workload scheduling across CPUs, GPUs, Kubernetes, bare metal, and edge devices.
  • Create reliable systems for checkpointing, retrying, resuming, scaling, and observing long-running workloads.
  • Develop data infrastructure for video, audio, sensor streams, telemetry, trajectories, policy traces, incidents, and interventions.
  • Design storage, indexing, lineage, replay, retention, and movement of large multimodal datasets.
  • Create APIs, SDKs, CLIs, and deployment workflows for ease of use by robotics engineers.
  • Own platform reliability, observability, capacity management, incident response, security, tenancy, and cost efficiency.
  • Support disconnected, bandwidth-constrained, private, air-gapped environments.
  • Connect cloud-side infra to edge runtime and operate as one platform.

Skills

Kubernetes
IaC
Networking
Storage
Security
Cloud architecture
GPU infrastructure
MLOps
Distributed systems
Rust/Go/Python

Tools

Ray
Kueue
Temporal
Kafka
NATS
ClickHouse
Parquet

Job description

The role:

Build the control plane and data plane for physical AI.

Munari must run training, simulation, evaluation, data processing, and inference workloads across public cloud, customer-owned infrastructure, robotics labs, and fleets of edge devices. It must handle GPUs, enormous multimodal datasets, unreliable connectivity, long-running workloads, and machines that cannot simply be restarted whenever something goes wrong.

This is not a conventional DevOps role and it is not a YAML-only role. You will write the systems software that makes physical AI infrastructure feel closer to a programmable platform than a pile of bespoke operations.

What you will work on:
  • Build a unified execution platform for simulation, policy evaluation, model training, data processing, and deployment.

  • Design workload scheduling and orchestration across CPUs, GPUs, Kubernetes clusters, bare metal, on-prem environments, and edge devices.

  • Build reliable systems for checkpointing, retrying, resuming, scaling, and observing long-running robotics workloads.

  • Create the data infrastructure for video, audio, sensor streams, telemetry, trajectories, policy traces, incidents, and human interventions.

  • Design storage, indexing, lineage, replay, retention, and movement of large multimodal datasets.

  • Build the APIs, SDKs, CLIs, and deployment workflows that make the underlying infrastructure simple for robotics engineers and researchers.

  • Own platform reliability, observability, capacity management, incident response, security, tenancy, and cost efficiency.

  • Support disconnected, bandwidth-constrained, private, and potentially air-gapped customer environments.

  • Connect cloud-side infrastructure to the Munari edge runtime and make the entire system operable as one platform.

You may be a strong fit if:
  • You have built or operated large-scale distributed systems where reliability genuinely mattered.

  • You have strong experience with Kubernetes, infrastructure as code, networking, storage, security, and cloud architecture.

  • You understand GPU infrastructure, distributed training, batch systems, MLOps, or high-performance computing.

  • You have worked with high-volume streaming, time-series, image, video, or scientific data.

  • You write production software in Rust, Go, Python, or a similar systems language rather than treating infrastructure as configuration alone.

  • You are comfortable debugging across an application, container, scheduler, network, host, GPU, and storage system.

  • You care equally about the internal architecture and the developer experience exposed to the customer.

Experience with NVIDIA infrastructure, Slurm, Ray, Kueue, Temporal, Kafka, NATS, ClickHouse, Parquet, object storage, multi-cluster Kubernetes, or edge fleet management is useful but not mandatory.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer (Edge Systems)
Founding Engineer (Edge Systems)

a16z-speedrun • San Francisco (CA)

On-site
USD 180,000 - 230,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Distributed Systems & AI Infrastructure Engineer
Distributed Systems & AI Infrastructure Engineer

Krea • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

descript • United States

On-site
USD 220,000 - 292,000
Equity
Competitive benefits
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000
Engineer, Supercomputing & Distributed Systems
Engineer, Supercomputing & Distributed Systems

Krea • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000