Turn this role into an interview — a resume and cover letter built around what this employer wants.
Build AI, an in-person team in San Francisco, seeks a skilled ML/systems engineer to optimize inference performance, reduce latency and cost, and enable scaling of the data engine. You will work closely with research and product to ensure affordable, accurate models and efficient serving.
Ideal candidates think in dollars and tokens per second, are proficient in Python and C++/Rust, and are comfortable profiling and tuning GPUs/accelerators in a small, cost–constrained team.
Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.
Inference is about 90% of compute spend. Economics are heavily driven by inference optimization. We’re hiring someone to make inference cheaper, faster, and good enough that we can scale the data engine and the product without the GPU bill eating the company.
Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.
We are a fully in‑person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application.
Questions: research@build.ai