Frontier LLM Inference Runtime Engineer

Triwill Group

United States

On-site

USD 180,000 - 240,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a highly skilled systems programmer to build the model runtime within the inference engine that runs frontier models on OpenAI’s custom silicon. You will bridge between models and the cluster serving software, translating workloads into efficient execution while optimizing throughput, latency, and reliability.

You will collaborate with model architecture, distributed systems, compilers, kernels, and silicon teams to co-design interfaces and remove bottlenecks.

Qualifications

  • Experience building or optimizing runtimes and distributed systems.
  • Familiarity with model-serving infrastructure and systems software.
  • Understanding of LLM inference, batching, KV-cache tradeoffs.

Responsibilities

  • Design and implement the LLM inference runtime for frontier models on custom silicon.
  • Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
  • Develop distributed execution strategies across chips, hosts, and racks.

Skills

C++
Rust
Python
Distributed systems
Model inference

Job description

OpenAI is seeking a highly skilled systems programmer to build the model runtime within the inference engine that runs frontier models on OpenAI’s custom silicon. You will bridge between models and the cluster serving software, translating workloads into efficient execution while optimizing throughput, latency, and reliability.

You will collaborate with model architecture, distributed systems, compilers, kernels, and silicon teams to co-design interfaces and remove bottlenecks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model Runtime Engineer for Frontier AI Inference
Model Runtime Engineer for Frontier AI Inference

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Runtime Engineer — LLM Inference & Host Stack
Runtime Engineer — LLM Inference & Host Stack

MatX Inc. • Mountain View (CA)

Hybrid
USD 160,000 - 475,000
Time off: 4 weeks PTO + 12 holidays +
Health: Company-subsidized Medical + D
Financial Wellbeing: 401K with company
+5