Stand out for this role — generate a tailored resume and cover letter in about a minute.
Acceler8 Talent is recruiting an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that is building an eight-figure revenue and growth trajectory. You will design, implement, and optimize the infrastructure powering large-scale LLM workloads and real-time model serving.
Ideal candidates combine Python/C++ proficiency with distributed systems experience, PyTorch expertise, and a passion for low-latency, GPU-accelerated inference at production scale.
Acceler8 Talent is recruiting an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that is building an eight-figure revenue and growth trajectory. You will design, implement, and optimize the infrastructure powering large-scale LLM workloads and real-time model serving.
Ideal candidates combine Python/C++ proficiency with distributed systems experience, PyTorch expertise, and a passion for low-latency, GPU-accelerated inference at production scale.