An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Acceler8 Talent is seeking a Member of Technical Staff to join its AI infrastructure team in San Francisco. You will design and build production-grade ML inference and model serving systems, optimizing latency, throughput, and resource utilization across large-scale AI workloads.
You will work with inference runtimes such as vLLM or TensorRT-LLM, tackle memory management and scheduling challenges, and collaborate with compiler, kernel, and distributed systems engineers to push the performance of
Acceler8 Talent is seeking a Member of Technical Staff to join its AI infrastructure team in San Francisco. You will design and build production-grade ML inference and model serving systems, optimizing latency, throughput, and resource utilization across large-scale AI workloads.
You will work with inference runtimes such as vLLM or TensorRT-LLM, tackle memory management and scheduling challenges, and collaborate with compiler, kernel, and distributed systems engineers to push the performance of