LLM Infrastructure Engineer — Scalable AI Systems

Get AI tool

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OpenAI is seeking engineers to design and build the core agent harness and execution loop for Codex agents in a production setting. You will craft sandboxing, isolation, and workflow infrastructure to support reliable agent orchestration and evaluation across multiple model interfaces.

You will run experiments on prompts, tool-use strategies, and runtime constraints, while improving observability and diagnostics for GPUs and inference systems. Strong collaboration with research is required.

Qualifications

  • Experience building or operating production distributed systems, infrastructure, tooling, sandboxing, or ML systems.
  • Ability to work across Rust systems code, Python configuration, APIs, and agent orchestration.
  • Hands-on experience with LLM applications, coding agents, model deployment, or inference systems.
  • Strong focus on reliability, safety, performance, and debuggability.
  • Experience collaborating with research while shipping production systems.
  • Strong coding skills and ownership of cross-functional AI projects.

Responsibilities

  • Design and build the core agent harness and execution loop for Codex agents.
  • Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents.
  • Develop evaluation, experimentation, and debugging systems across the agent stack.
  • Run experiments on prompts, model interfaces, context construction, and tool-use strategies.
  • Improve observability, profiling, and diagnostics across backend systems and GPUs.
  • Collaborate with research to make the harness trainable and measurable.
  • Build shared platform primitives to improve Codex speed and safety.

Skills

Distributed systems
Rust systems code
Python configuration
Agent orchestration
LLM applications
Reliability and debugging
Research collaboration
Platform engineering

Tools

Rust
Python
Linux
Sandboxing
Cloud platforms
GPU systems

Job description

OpenAI is seeking engineers to design and build the core agent harness and execution loop for Codex agents in a production setting. You will craft sandboxing, isolation, and workflow infrastructure to support reliable agent orchestration and evaluation across multiple model interfaces.

You will run experiments on prompts, tool-use strategies, and runtime constraints, while improving observability and diagnostics for GPUs and inference systems. Strong collaboration with research is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Agents Infrastructure Engineer
AI Agents Infrastructure Engineer

OpenAI • San Francisco (CA)

On-site
USD 230,000 - 385,000
Agent Harness Engineer: Build Safe, Scalable AI
Agent Harness Engineer: Build Safe, Scalable AI

OpenAI • San Francisco (CA)

On-site
USD 230,000 - 385,000
AI Systems Engineer, Codex Agents
AI Systems Engineer, Codex Agents

Get AI tool • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
AI Systems Engineer, Codex Agents
AI Systems Engineer, Codex Agents

OpenAI • San Francisco (CA)

On-site
USD 230,000 - 385,000
Software Engineer, Codex Core Agents
Software Engineer, Codex Core Agents

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000
Applied AI Engineer, Codex Core Agent
Applied AI Engineer, Codex Core Agent

Slope • San Francisco (CA)

On-site
USD 100,000 - 140,000
AI Systems Engineer, Codex Agents
AI Systems Engineer, Codex Agents

Slope • San Francisco (CA)

On-site
USD 230,000 - 385,000
Medical, dental, and vision insurance
401(k) retirement plan with employer match
Paid parental leave
+2
Engineer, Codex Core Agents — Scalable AI Agent Infra
Engineer, Codex Core Agents — Scalable AI Agent Infra

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000
Staff Software Engineer, AI Infrastructure
Staff Software Engineer, AI Infrastructure

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation package
Significant technical ownership
Meaningful equity
Remote AI/ML Engineer: Scalable LLMs & Multi-Agent Apps
Remote AI/ML Engineer: Scalable LLMs & Multi-Agent Apps

YO AI Labs • North Carolina

Remote
USD 120,000 - 160,000