Get more replies from employers
Send a job-specific resume in minutes.
Build AI is hiring a lead for the data platform to manage on-device capture, upload under flaky bandwidth, and training-ready shards for research customers. The role spans compression, storage-tier decisions, and end-to-end dataset packaging and delivery.
You will collaborate with Shenzhen firmware to standardize ingest contracts and help scale across sites worldwide. The team values engineers who can work across engineering and research, shipping data paths at petabyte scale and beyond in a
Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.
We’re hiring a lead for the data platform: camera on a worker to training-ready datasets, and out to research customers. Collection is monocular 1920×1080p 30fps in the wild, targeting 100M hours. The roles we mean: Tesla Autopilot, Waymo, Cruise, Zoox, Nuro, Samsara, Verkada, Netflix encoding, YouTube ingest, Scale, Eventual/Daft. This is not a warehouse, analytics, or generic backend seat.
Own the data platform end-to-end: on-device capture, upload under flaky bandwidth, object storage, training-ready shards
Compression, codecs, and storage-tier trade-offs so 1080p30 hours stay cheap enough to keep collecting
Upload that survives bad networks: on-device buffering, batching, retries, a drop rate you can actually see
Object storage and training-shard formats. The hard problem is petabyte-scale media, not a warehouse
Own dataset packaging, versioning, and delivery to external research customers
Work with Shenzhen firmware so new devices speak one ingest contract, not a custom path per SKU
Make health, cost, and drop rate obvious as we add sites and countries
You have owned a production media or sensor data path at real scale: object storage at petabyte scale, video codecs and compression, upload under flaky bandwidth, or training-shard / dataset formats
That kind of data path: Tesla Autopilot, Waymo, Cruise, Zoox, Nuro, Samsara, Verkada, Netflix encoding, YouTube ingest, Scale, or Eventual/Daft. Demo-scale ETL is not this job
Strong software engineering. Python and at least one systems language. Linux
You measure cost and throughput, not whether the demo uploaded
You want to scale in-the-wild physical-labor video, not run a generic data org
Pose, multi-camera, or other large media besides video
Cloud (AWS or GCP), orchestration (Kubernetes, Airflow, Temporal), or IaC
Dataset management or annotation tooling
You have shipped dataset delivery to external research or training customers
Competitive pay
Medical, dental, and vision packages with generous premium coverage
$500 per month credit for waiving medical benefits
Housing subsidy of $2k per month for those living within walking distance of the office
Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)
Various wellness benefits covering fitness, mental health, and more
Daily lunch and dinner in our office
Unlimited compute budget subject to ROI justification
Unlimited Codex and Claude credits
Travel
Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.
We are a fully in‑person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: research@build.ai