Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
a16z-speedrun is seeking an experienced systems engineer to build and operate the control and data planes for physical AI across public clouds, customer-owned infrastructure, robotics labs, and edge fleets. You will design orchestration across CPUs/GPUs, Kubernetes, bare metal, on-prem, and edge devices, while ensuring reliability for long-running workloads and complex data pipelines.
Role emphasizes production-grade Rust/Go/Python software, strong debugging across app, container, scheduler,
Build the control plane and data plane for physical AI.
Munari must run training, simulation, evaluation, data processing, and inference workloads across public cloud, customer-owned infrastructure, robotics labs, and fleets of edge devices. It must handle GPUs, enormous multimodal datasets, unreliable connectivity, long-running workloads, and machines that cannot simply be restarted whenever something goes wrong.
This is not a conventional DevOps role and it is not a YAML-only role. You will write the systems software that makes physical AI infrastructure feel closer to a programmable platform than a pile of bespoke operations.
Build a unified execution platform for simulation, policy evaluation, model training, data processing, and deployment.
Design workload scheduling and orchestration across CPUs, GPUs, Kubernetes clusters, bare metal, on-prem environments, and edge devices.
Build reliable systems for checkpointing, retrying, resuming, scaling, and observing long-running robotics workloads.
Create the data infrastructure for video, audio, sensor streams, telemetry, trajectories, policy traces, incidents, and human interventions.
Design storage, indexing, lineage, replay, retention, and movement of large multimodal datasets.
Build the APIs, SDKs, CLIs, and deployment workflows that make the underlying infrastructure simple for robotics engineers and researchers.
Own platform reliability, observability, capacity management, incident response, security, tenancy, and cost efficiency.
Support disconnected, bandwidth-constrained, private, and potentially air-gapped customer environments.
Connect cloud-side infrastructure to the Munari edge runtime and make the entire system operable as one platform.
You have built or operated large-scale distributed systems where reliability genuinely mattered.
You have strong experience with Kubernetes, infrastructure as code, networking, storage, security, and cloud architecture.
You understand GPU infrastructure, distributed training, batch systems, MLOps, or high-performance computing.
You have worked with high-volume streaming, time-series, image, video, or scientific data.
You write production software in Rust, Go, Python, or a similar systems language rather than treating infrastructure as configuration alone.
You are comfortable debugging across an application, container, scheduler, network, host, GPU, and storage system.
You care equally about the internal architecture and the developer experience exposed to the customer.
Experience with NVIDIA infrastructure, Slurm, Ray, Kueue, Temporal, Kafka, NATS, ClickHouse, Parquet, object storage, multi-cluster Kubernetes, or edge fleet management is useful but not mandatory.