Get more replies from employers
Send a job-specific resume in minutes.
Jobtailor in Palo Alto is seeking an experienced ML inference engineer to build and optimize high-throughput inference services for large multimodal models. You will implement multi-GPU and multi-node parallelism, tune latency and throughput, and collaborate with researchers to push new architectures.
This on-site role emphasizes reliability, observability, and efficiency, with opportunities to shape production ML stacks and optimize GPU fleets at scale.