A complete application in a minute — tailored resume and cover letter, ready to send.
Build AI in San Francisco seeks a lead for the data platform, overseeing on-device capture, upload under flaky bandwidth, and training-ready dataset pipelines. You will design storage, formats, and tooling to support petabyte-scale media for researchers and clients, including coordination with Shenzhen firmware teams.
You will optimize codecs, compression, and storage costs to keep 1080p30 hours affordable while scaling across sites and countries, balancing performance with ROI.
Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.
We’re hiring a lead for the data platform: camera on a worker to training-ready datasets, and out to research customers. Collection is monocular 1920×1080p 30fps in the wild, targeting 100M hours. The roles we mean: Tesla Autopilot, Waymo, Cruise, Zoox, Nuro, Samsara, Verkada, Netflix encoding, YouTube ingest, Scale, Eventual/Daft. This is not a warehouse, analytics, or generic backend seat.
Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.
We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application. Questions: research@build.ai