Distributed Runtime Foundations Engineer

OpenAI

San Francisco (CA)

Hybrid

USD 220,000 - 300,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Relocation assistance

Job summary

OpenAI is seeking a Training Runtime: Runtime Foundations Engineer in San Francisco, CA to help tie thousands of computers into a unified, high-performance system. You will work primarily in Codex + Rust, building asynchronous components to orchestrate and monitor ML workloads on our largest clusters.

You will optimize for scale, reliability, and observability, and collaborate across Python and Rust stacks while diving into novel domains and demanding performance challenges.

Qualifications

  • Experience developing distributed systems.
  • Strong understanding of large-scale system behavior and failure modes.
  • Proficiency in Rust or another systems language.
  • Solid Linux knowledge and memory profiling.
  • Experience with asynchronous and concurrent systems.

Responsibilities

  • Work across our Python and Rust stack.
  • Design, build, and maintain software to orchestrate and monitor machine learning workloads on our largest supercomputers.
  • Profile and optimize our software stack to support computation orchestration at frontier scale.
  • Improve reliability, observability, and fault tolerance for long-running jobs.
  • Debug complex distributed systems issues across large clusters.
  • Respond to the changing shapes and needs of the ML systems to enable our researchers.

Skills

Distributed systems
Rust
C/C++
Linux
Asynchronous programming
Performance optimization

Job description

OpenAI is seeking a Training Runtime: Runtime Foundations Engineer in San Francisco, CA to help tie thousands of computers into a unified, high-performance system. You will work primarily in Codex + Rust, building asynchronous components to orchestrate and monitor ML workloads on our largest clusters.

You will optimize for scale, reliability, and observability, and collaborate across Python and Rust stacks while diving into novel domains and demanding performance challenges.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Training - Runtime Foundations Engineer
Training - Runtime Foundations Engineer

OpenAI • San Francisco (CA)

On-site
USD 220,000 - 300,000
Relocation assistance
Training Performance Engineer
Training Performance Engineer

Slope • San Francisco (CA)

On-site
USD 250,000 - 460,000
Relocation assistance
Flexible working hours
Collaborative work environment
Training: ML Framework Engineer
Training: ML Framework Engineer

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Relocation assistance
Hybrid work model
Model Runtime Engineer for Frontier AI Inference
Model Runtime Engineer for Frontier AI Inference

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Software Engineer, Platform Systems
Software Engineer, Platform Systems

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
AI Agents Infrastructure Engineer
AI Agents Infrastructure Engineer

OpenAI • San Francisco (CA)

On-site
USD 230,000 - 385,000
Model Runtime Engineer for Frontier AI on Custom Silicon
Model Runtime Engineer for Frontier AI on Custom Silicon

OpenAI, Inc. • San Francisco (CA)

On-site
USD 266,000 - 445,000
Equity
Health insurance
401(k) match
+2
Compute Foundations Engineer — Kubernetes & GPU Infra
Compute Foundations Engineer — Kubernetes & GPU Infra

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 490,000
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
Model Runtime Engineer for Frontier AI on Custom Silicon
Model Runtime Engineer for Frontier AI on Custom Silicon

OpenAI • United States

Remote
USD 180,000 - 230,000