Model Runtime Engineer for Frontier AI on Custom Silicon

OpenAI

United States

Remote

USD 180,000 - 230,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OpenAI is seeking a skilled engineer to build the model runtime within the inference engine that executes frontier models on our custom silicon. You will bridge models on hardware with the cluster serving software, translating workloads into efficient execution while optimizing throughput, latency, and reliability.

You will collaborate across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime, comparable to systems like vLLM but tailored

Responsibilities

  • Design and implement the LLM inference runtime for frontier models running on custom silicon.
  • Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
  • Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
  • Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads.
  • Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack.

Job description

OpenAI is seeking a skilled engineer to build the model runtime within the inference engine that executes frontier models on our custom silicon. You will bridge models on hardware with the cluster serving software, translating workloads into efficient execution while optimizing throughput, latency, and reliability.

You will collaborate across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime, comparable to systems like vLLM but tailored

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Runtime Engineer for Frontier AI on Custom Silicon
Model Runtime Engineer for Frontier AI on Custom Silicon

OpenAI, Inc. • San Francisco (CA)

On-site
USD 266,000 - 445,000
Equity
Health insurance
401(k) match
+2
Model Runtime Engineer for Frontier AI Inference
Model Runtime Engineer for Frontier AI Inference

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Frontier LLM Inference Runtime Engineer
Frontier LLM Inference Runtime Engineer

Triwill Group • United States

On-site
USD 180,000 - 240,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • United States

Remote
USD 180,000 - 230,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 250,000 - 360,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000
CI/CD Architect for Frontier AI Hardware
CI/CD Architect for Frontier AI Hardware

OpenAI • San Francisco (CA)

On-site
USD 177,000 - 327,000
Low-Level Runtime Engineer for AI Accelerator Silicon
Low-Level Runtime Engineer for AI Accelerator Silicon

OpenAI • United States

Remote
USD 160,000 - 220,000