An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Inferact is seeking an inference runtime engineer to enhance the performance and capabilities of LLM and diffusion model serving. This role requires expertise in optimizing model execution on various hardware architectures and has significant implications for AI inference.
The ideal candidate must possess a bachelor's degree in computer science or related fields, strong programming skills in Python, and experience with LLM inference systems. Remote work options are available for exceptional candidates.
We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.
Minimum qualifications:
Preferred qualifications:
Bonus points if you have: