An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner teams to integrate external models.
The role requires deep expertise in PyTorch, TensorRT, CUDA, and model optimization techniques, with a focus on delivering cutting-edge inference capabilities at scale.
We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models.
You'll work across the inference stack, designing novel frameworks, optimizing inference performance, and shaping Reactor's competitive edge in ultra-low-latency, high-throughput environments.
Engineering
San Francisco