Turn this role into an interview — a resume and cover letter built around what this employer wants.
Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, and work at the core of vLLM to accelerate AI inference.
The role requires deep knowledge of transformer models, strong Python and PyTorch skills, and experience with LLM inference systems. You will read papers, implement techniques, and contribute robust, maintainable code to complex ML systems.
Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, and work at the core of vLLM to accelerate AI inference.
The role requires deep knowledge of transformer models, strong Python and PyTorch skills, and experience with LLM inference systems. You will read papers, implement techniques, and contribute robust, maintainable code to complex ML systems.