An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Acceler8 Talent is recruiting an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that is building an eight-figure revenue and growth trajectory. You will design, implement, and optimize the infrastructure powering large-scale LLM workloads and real-time model serving.
Ideal candidates combine Python/C++ proficiency with distributed systems experience, PyTorch expertise, and a passion for low-latency, GPU-accelerated inference at production scale.
We’re hiring anML Inference Engineerfor a Stanford-spun AI startup in San Francisco that has already grown to8-figure revenue.
The team is rebuilding itsLLM inference stack from the ground up, solving challenging systems problems around GPU performance, distributed compute and real-time model serving.
This role is for engineers who enjoy going deep onperformance, infrastructure, and optimisation.