Production-Grade LLM Inference Runtime Engineer

AI Chopping Block

San Francisco, Northern (CA, KY)

Hybrid

USD 250,000 - 360,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a systems-focused engineer to design and implement the LLM inference runtime for frontier models on our custom silicon. You'll bridge model execution with the hardware, shaping how workloads map onto the platform and how production workloads achieve high throughput and low latency.

You will collaborate across model, systems, compiler, kernel, and hardware teams to build a scalable, observable runtime and to deliver reliable, production-grade performance on OpenAI's AI

Qualifications

  • Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
  • Experience building or optimizing runtimes, distributed systems, or model-serving infrastructure.
  • Understanding of modern LLM inference including batching, prefill/decode behavior, and model parallelism.
  • Ability to reason about latency, throughput, memory bandwidth, and utilization.
  • Experience profiling and debugging across hardware-software stacks.
  • Able to design clean abstractions with low-level control to extract performance from specialized hardware.

Responsibilities

  • Design and implement the LLM inference runtime for frontier models running on custom silicon.
  • Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
  • Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
  • Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads.
  • Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack.
  • Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
  • Create profiling, observability, benchmarking, and performance-modeling tools that make runtime behavior measurable and actionable.
  • Debug complex correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware.
  • Turn workload insights into clear requirements for future generations of silicon and system architecture.

Skills

C++
Rust
Python
Distributed systems
Systems programming
Performance optimization

Job description

OpenAI is seeking a systems-focused engineer to design and implement the LLM inference runtime for frontier models on our custom silicon. You'll bridge model execution with the hardware, shaping how workloads map onto the platform and how production workloads achieve high throughput and low latency.

You will collaborate across model, systems, compiler, kernel, and hardware teams to build a scalable, observable runtime and to deliver reliable, production-grade performance on OpenAI's AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier LLM Inference Runtime Engineer
Frontier LLM Inference Runtime Engineer

Triwill Group • United States

On-site
USD 180,000 - 240,000
LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
Model Runtime Engineer for Frontier AI Inference
Model Runtime Engineer for Frontier AI Inference

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA