Stand out for this role — generate a tailored resume and cover letter in about a minute.
Inferact is seeking an inference runtime engineer to advance LLM and diffusion model serving. You will optimize how models execute across diverse hardware, shaping the core of vLLM and enabling faster AI inference.
This remote role embraces flexible timezones with Pacific overlap for critical syncs, and compensation includes salary plus equity. The ideal candidate will have deep knowledge of transformer models, strong Python/PyTorch skills, and hands-on experience with LLM inference systems.
Inferact is seeking an inference runtime engineer to advance LLM and diffusion model serving. You will optimize how models execute across diverse hardware, shaping the core of vLLM and enabling faster AI inference.
This remote role embraces flexible timezones with Pacific overlap for critical syncs, and compensation includes salary plus equity. The ideal candidate will have deep knowledge of transformer models, strong Python/PyTorch skills, and hands-on experience with LLM inference systems.