Get more replies from employers
Send a job-specific resume in minutes.
Build AI is hiring to cut inference costs and improve latency and throughput for scalable physical AI workloads. You will optimize serving, quantization, and kernel performance while collaborating with research and product teams to keep models accurate yet affordable at scale.
The role focuses on profiling, architectural tweaks, and cost-aware engineering in a small, cost-conscious team in a fully in-person setup across San Francisco and Shenzhen.
Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.
Inference is about 90% of compute spend. Economics are heavily driven by inference optimization. We're hiring someone to make inference cheaper, faster, and good enough that we can scale the data engine and the product without the GPU bill eating the company.
Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.
We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: research@build.ai