Software Engineer, Inference Runtime

LM Studio

New York (NY)

Hybrid

USD 150,000 - 350,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity grants
Great medical, vision, dental plans
Catered team lunch / expensed dinners
Flexible PTO
Flexible WFH
Sun-drenched office in SoHo in NYC

Job summary

LM Studio is seeking an Inference Runtime Software Engineer to push forward our on-device and cloud inference stack. You will integrate new engines, bring up open-weight models, and optimize execution across CPU/GPU targets.

Join a high-intensity team building human-centered AI tools, contributing to open-source projects like llama.cpp, MLX, and vLLM, with competitive compensation and flexible work options in NYC.

Qualifications

  • Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure.
  • Strong programming ability in Python and C++.
  • Deep understanding of transformer architectures and the mechanics of model inference.
  • Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement.
  • Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM.
  • Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution.
  • Takes personal responsibility for the correctness and performance of their work.

Responsibilities

  • Maintain and push forward our inference stack on-device and in the cloud.
  • Bring up new model architectures and multimodal models.
  • Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes.
  • Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution.
  • Benchmark and diagnose correctness and performance problems across the inference stack.
  • Contribute upstream to open-source projects such as llama.cpp and MLX.

Skills

Python
C++
Transformer models
Profiling
PyTorch
Inference systems
Debugging
Ownership

Tools

llama.cpp
MLX
ExecuTorch
vLLM
SGLang
TensorRT-LLM

Job description

LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family.


As a team, we work with high technical intensity and personal responsibility. We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software.


The Role

We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on.


Qualifications


  • Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure

  • Strong programming ability in Python and C++

  • Deep understanding of transformer architectures and the mechanics of model inference

  • Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement

  • Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

  • Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution

  • Takes personal responsibility for the correctness and performance of their work


Bonus Qualifications


  • Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM


Responsibilities


  • Maintain and push forward our inference stack on-device and in the cloud

  • Bring up new model architectures and multimodal models

  • Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes

  • Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution

  • Benchmark and diagnose correctness and performance problems across the inference stack
  • Contribute upstream to open-source projects such as llama.cpp and MLX


Benefits


  • Competitive salary and equity grants

  • Great medical, vision, dental healthcare plans

  • Catered team lunch / expensed dinners in the office

  • Flexible PTO

  • Flexible WFH

  • Sun-drenched office in SoHo in NYC


Compensation Range: $150K - $350K

Get your free, confidential resume review.
or drag and drop your file here.