Turn this role into an interview — a resume and cover letter built around what this employer wants.
Acceler8 Talent in San Francisco seeks an ML Inference Engineer to redesign the infrastructure behind its production AI platform. You will architect high-performance systems for serving LLMs, optimize inference for latency and throughput, and scale GPU workloads across a Kubernetes-based stack.
You will work close to the hardware and software stack, balancing trade-offs for impactful improvements at scale.
A Stanford-spun AI company in San Francisco is hiring an ML Inference Engineer to help redesign the infrastructure behind its production AI platform.
What you'll be solving
What you'll bring
This is an opportunity to own meaningful parts of an inference stack being built from the ground up, rather than maintaining an established platform.