An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Acceler8 Talent in San Francisco seeks an ML Inference Engineer to redesign the infrastructure behind its production AI platform. You will architect high-performance systems for serving LLMs, optimize inference for latency and throughput, and scale GPU workloads across a Kubernetes-based stack.
You will work close to the hardware and software stack, balancing trade-offs for impactful improvements at scale.
Acceler8 Talent in San Francisco seeks an ML Inference Engineer to redesign the infrastructure behind its production AI platform. You will architect high-performance systems for serving LLMs, optimize inference for latency and throughput, and scale GPU workloads across a Kubernetes-based stack.
You will work close to the hardware and software stack, balancing trade-offs for impactful improvements at scale.