Get more replies from employers
Send a job-specific resume in minutes.
LM Studio in New York is seeking an Inference Runtime Software Engineer to advance our on-device and cloud inference stack. You will integrate new inference engines, optimize execution across CPU and GPU, and contribute to open-source projects.
You will also bring up new models and modalities, improve latency and memory use, and ensure reliability across diverse runtimes and hardware. Join a highly skilled team shaping human‑AI interactions.
LM Studio in New York is seeking an Inference Runtime Software Engineer to advance our on-device and cloud inference stack. You will integrate new inference engines, optimize execution across CPU and GPU, and contribute to open-source projects.
You will also bring up new models and modalities, improve latency and memory use, and ensure reliability across diverse runtimes and hardware. Join a highly skilled team shaping human‑AI interactions.