Runtime Engineer — LLM Inference & Host Stack

MatX Inc.

Mountain View (CA)

Hybrid

USD 160,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Time off: 4 weeks PTO + 12 holidays +
Health: Company-subsidized Medical + D
Financial Wellbeing: 401K with company
Professional Development Budget
Team meals and commuting reimbursement
AI resources assistance
Mental wellbeing benefits
Parental leave

Job summary

MatX Inc. seeks a systems programmer to build the host-side interface library and manage the compiler→runtime contract. You will design the custom-kernel ABI, and implement Python bindings to move tensors from Python to accelerator hardware.

The role involves working with CUDA/ROCm-style accelerators, memory models, and high-performance computing stacks, delivering efficient runtime and serving throughput.

Qualifications

  • Experience in a systems programming language with memory management and ABI work.
  • Production Python interop layers (PyO3, ctypes, pybind11) experience.
  • Experience designing and maintaining API/ABI contracts between teams.

Responsibilities

  • Build the host-side interface library: memory management, DMA, streams and events, and sync primitives.
  • Own and extend the executable format: compiler→runtime contract, versioning, weight/quantization layouts.
  • Design the custom-kernel ABI and host-side marshaling for Python tensors to device.

Skills

Rust
C/C++
Go
FFI/ABI
Python interop

Job description

MatX Inc. seeks a systems programmer to build the host-side interface library and manage the compiler→runtime contract. You will design the custom-kernel ABI, and implement Python bindings to move tensors from Python to accelerator hardware.

The role involves working with CUDA/ROCm-style accelerators, memory models, and high-performance computing stacks, delivering efficient runtime and serving throughput.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Runtime Systems Engineer - Host Stack & AI Inference
Runtime Systems Engineer - Host Stack & AI Inference

MatX • Mountain View (CA)

On-site
USD 160,000 - 475,000
Time off
Health benefits
Retirement plan
+6
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Compiler Runtime Engineer
Compiler Runtime Engineer

Oho Group • San Francisco (CA)

On-site
USD 150,000 - 210,000
Frontier LLM Inference Runtime Engineer
Frontier LLM Inference Runtime Engineer

Triwill Group • United States

On-site
USD 180,000 - 240,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
Runtime Engineer
Runtime Engineer

Lemurian Labs Inc. • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity
Medical/Dental/Vision
Retirement savings plan
+1
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Runtime Engineer
Runtime Engineer

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity opportunities
Medical, dental, and vision coverage
Retirement savings plan
+2
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000