Get more replies from employers
Send a job-specific resume in minutes.
LM Studio is seeking an Inference Runtime Software Engineer to push forward our on-device and cloud inference stack. You will integrate new engines, bring up open-weight models, and optimize execution across CPU/GPU targets.
Join a high-intensity team building human-centered AI tools, contributing to open-source projects like llama.cpp, MLX, and vLLM, with competitive compensation and flexible work options in NYC.
LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family.
As a team, we work with high technical intensity and personal responsibility. We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software.
We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on.
Compensation Range: $150K - $350K