Model Runtime Engineer for Frontier AI on Custom Silicon

OpenAI, Inc.

San Francisco (CA)

On-site

USD 266,000 - 445,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
401(k) match
Paid parental leave
Flexible PTO

Job summary

OpenAI, Inc. is seeking a Software Engineer to build the model runtime for frontier models on OpenAI's custom silicon in San Francisco.

You will optimize inference throughput, latency, and reliability within the inference engine and coordinate with kernel, compiler, architecture, and silicon teams to remove bottlenecks across the stack. The role requires strong systems programming in C++, Rust, Python with experience in runtimes or model-serving infrastructure, and a capability to design

Qualifications

  • Strong systems programming experience in C++, Rust, Python or comparable performance oriented environments.
  • Built or optimized runtimes, distributed systems, compilers, kernels, or model-serving infrastructure.

Responsibilities

  • Design and implement the LLM inference runtime for frontier models running on custom silicon.
  • Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
  • Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
  • Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads.
  • Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack.
  • Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
  • Create profiling, observability, benchmarking, and performance-modeling tools that make runtime behavior measurable and actionable.
  • Debug complex correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware.
  • Turn workload insights into clear requirements for future generations of silicon and system architecture.

Skills

C++
Rust
Python
Distributed systems

Job description

OpenAI, Inc. is seeking a Software Engineer to build the model runtime for frontier models on OpenAI's custom silicon in San Francisco.

You will optimize inference throughput, latency, and reliability within the inference engine and coordinate with kernel, compiler, architecture, and silicon teams to remove bottlenecks across the stack. The role requires strong systems programming in C++, Rust, Python with experience in runtimes or model-serving infrastructure, and a capability to design

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Runtime Engineer for Frontier AI on Custom Silicon
Model Runtime Engineer for Frontier AI on Custom Silicon

OpenAI • United States

Remote
USD 180,000 - 230,000
Model Runtime Engineer for Frontier AI Inference
Model Runtime Engineer for Frontier AI Inference

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Frontier LLM Inference Runtime Engineer
Frontier LLM Inference Runtime Engineer

Triwill Group • United States

On-site
USD 180,000 - 240,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • United States

Remote
USD 180,000 - 230,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 250,000 - 360,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000
CI/CD Architect for Frontier AI Hardware
CI/CD Architect for Frontier AI Hardware

OpenAI • San Francisco (CA)

On-site
USD 177,000 - 327,000
Low-Level Runtime Engineer for AI Accelerator Silicon
Low-Level Runtime Engineer for AI Accelerator Silicon

OpenAI • United States

Remote
USD 160,000 - 220,000
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000