Model Runtime Engineer for Frontier AI Inference

OpenAI

San Francisco (CA)

On-site

USD 266,000 - 445,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a seasoned systems programmer to build the model runtime for frontier models on our AI accelerator. You will design an inference runtime that sits between model execution on hardware and the serving stack, optimizing throughput and latency.

Collaborating across model architecture, distributed systems, compilers and kernel teams, you will push toward production-grade reliability, observability, and scalable performance for OpenAI's supercomputing platform.

Qualifications

  • Strong systems programming experience in C++, Rust, or Python.
  • Experience building or optimizing runtimes, distributed systems, or compilers.
  • Understand LLM inference, including prefill, decode, batching, and model parallelism.
  • Ability to reason about latency, throughput, memory bandwidth, and utilization.

Responsibilities

  • Design and implement the LLM inference runtime for frontier models on custom silicon.
  • Build scheduling, memory management, and execution orchestration for high-performance inference.
  • Develop distributed execution strategies across chips, hosts, and racks.
  • Optimize end-to-end latency, throughput, and hardware utilization.
  • Create profiling, observability, and performance-modeling tools.
  • Debug correctness and reliability issues across software and hardware.

Skills

C++
Rust
Python
Distributed systems
Systems programming

Job description

OpenAI is seeking a seasoned systems programmer to build the model runtime for frontier models on our AI accelerator. You will design an inference runtime that sits between model execution on hardware and the serving stack, optimizing throughput and latency.

Collaborating across model architecture, distributed systems, compilers and kernel teams, you will push toward production-grade reliability, observability, and scalable performance for OpenAI's supercomputing platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier LLM Inference Runtime Engineer
Frontier LLM Inference Runtime Engineer

Triwill Group • United States

On-site
USD 180,000 - 240,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000
Platform Engineer for AI Inference & Optimization
Platform Engineer for AI Inference & Optimization

OpenAI • Seattle (WA)

On-site
USD 180,000 - 240,000
Software Engineer, AI Inference Infrastructure Platform
Software Engineer, AI Inference Infrastructure Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Accelerator Runtime Systems Engineer
AI Accelerator Runtime Systems Engineer

Triwill Group • United States

On-site
USD 180,000 - 240,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2