Get more replies from employers
Send a job-specific resume in minutes.
LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures, squeezing maximum performance while integrating with model-serving infrastructure.
The role emphasizes deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure to frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup building cutting-edge AI software.
Experience: Senior Level
Salary: $175,000 - $250,000 per year
We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing maximum performance from complex compute environments.
We're a well-funded AI infrastructure startup developing modern software at the intersection of artificial intelligence, high-performance computing, and specialized hardware. The team is tackling complex performance challenges associated with running modern AI workloads across emerging compute architectures.
#techservices #c #python #gpu #rust #dataflow #optimization #library #cuda #algorithm #latency #itl #tvm #llvm #multimodal #quantization #vllm #kv-cache #kernel-variants #ml-inference #systolic-arrays #ttft #tpot #tier3