Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, and work at the core of vLLM to accelerate AI inference.
The role requires deep knowledge of transformer models, strong Python and PyTorch skills, and experience with LLM inference systems. You will read papers, implement techniques, and contribute robust, maintainable code to complex ML systems.
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.
Minimum qualifications:
Preferred qualifications:
Bonus points if you have: