Principal Software Engineer, Simulation

Slope

San Francisco (CA)

On-site

USD 210,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI, SF-based, seeks a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You will own essential interfaces between research and engineering, building systems, APIs, and patterns enabling researchers to run agentic coding in training environments safely and efficiently.

The role suits senior backend/infrastructure engineers with strong judgment, product sense for technical users, and the ability to drive multi-team execution while ensuring

Qualifications

  • Significant experience building and scaling backend or infrastructure systems in fast-moving environments.
  • Deep strength in API design, systems design, and engineering fundamentals.
  • Detail-oriented with focus on correctness, reliability, and operational quality.
  • Ability to work with demanding technical users while maintaining engineering discipline.
  • Proven track record leading cross-functional technical efforts and driving clarity across teams.
  • Strong product sense and user empathy for internal platforms and developer tooling.
  • Proficiency in Python; backend platform engineering; Rust is a plus.

Responsibilities

  • Design, build, and evolve the integration between research harness and training infrastructure.
  • Build a platform for training and evaluating LLMs in simulated environments that match deployment settings.
  • Own major integration surfaces end-to-end, from architecture and API design through rollout and maintenance.
  • Build reliable execution systems that scale to demanding training workloads.
  • Partner with research, engineering, and platform teams to support new training use cases.
  • Design clean, stable interfaces and workflows for highly technical internal users.
  • Prevent short-term workarounds by establishing durable abstractions and clear ownership.
  • Raise the bar for correctness, reliability, and operational rigor in critical systems.

Skills

API design
Systems design
Backend infrastructure
Python
Rust
Cross-functional leadership
Technical judgment

Education

Bachelor's degree in CS or related field

Job description

About the Team

OpenAI’s research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI’s products and the research infrastructure where GPT‑next is trained, ensuring that our model’s training environment is as realistic as possible.

This team owns the integration layer that connects our production harness capabilities into the training stack. The work is highly cross‑functional and high leverage: researchers depend on it to run experiments and evaluations reliably as well as to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness.

About the Role

We’re looking for a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You’ll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments.

This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest‑leverage work is building robust infrastructure that supports and accelerates research without compromising engineering quality.

In this role, you will
  • Design, build, and evolve the integration between the Codex harness that powers OpenAI’s products and research training infrastructure used for training GPT‑next.

  • Build a platform for our LLMs to train and be evaluated in simulated environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more.

  • Own major integration surfaces end‑to‑end, from architecture and API design through rollout, operations, and long‑term maintenance.

  • Build reliable execution systems that can support demanding training workloads at scale.

  • Partner closely with research, agent, infrastructure, and platform teams to support new training use cases and harness capabilities.

  • Design clean, stable interfaces and workflows for highly technical internal users who move quickly and expect strong ergonomics.

  • Prevent one‑off workarounds from becoming long‑term technical debt by establishing durable abstractions and clear ownership.

  • Raise the bar for correctness, reliability, operational rigor, and engineering judgment across a critical research‑facing system.

You might thrive in this role if you
  • Have significant experience building and scaling backend or infrastructure systems in fast‑moving environments.

  • Bring deep strength in API design, systems design, and engineering fundamentals.

  • Are highly detail‑oriented and care deeply about correctness, reliability, and operational quality.

  • Can work directly with demanding technical users while maintaining strong engineering discipline.

  • Have a track record of leading cross‑functional technical efforts and creating clarity across organizational boundaries.

  • Bring strong product sense and user empathy for internal platforms and developer tooling.

  • Are motivated by enabling researchers and accelerating their work, rather than doing research yourself.

  • Are proficient in Python and have experience with backend platform engineering; Rust experience is a plus.

Location

This role is ideally based in San Francisco due to the close collaboration required with researchers and applied engineering partners.

We are an equal‑opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s affirmative action and equal employment opportunity policy statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US‑based candidates.

To notify OpenAI that you believe this job posting is non‑compliant, please submit a report through this form.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Simulation
Principal Software Engineer, Simulation

OpenAI • San Francisco (CA)

On-site
USD 347,000 - 490,000
Principal Software Engineer, Agent Harness Bridge
Principal Software Engineer, Agent Harness Bridge

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, API Agents
Software Engineer, API Agents

OpenAI • San Francisco (CA)

On-site
USD 170,000 - 230,000
Simulation Infrastructure Engineer
Simulation Infrastructure Engineer

OpenAI • San Francisco (CA)

On-site
USD 230,000 - 385,000
Software Engineer, ML Systems & Training Architecture
Software Engineer, ML Systems & Training Architecture

Neura Market • San Francisco (CA)

On-site
USD 295,000 - 380,000
Relocation assistance
Software Engineer, Cloud Agents
Software Engineer, Cloud Agents

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 385,000
Software Engineer, API Agents
Software Engineer, API Agents

Slope • San Francisco (CA)

On-site
USD 220,000 - 280,000
Software Engineer, RL Training Infra
Software Engineer, RL Training Infra

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer, ChatGPT Infrastructure
Software Engineer, ChatGPT Infrastructure

OpenAI • Los Angeles (CA)

Hybrid
USD 255,000 - 405,000
Software Engineer, Research - Human Data
Software Engineer, Research - Human Data

Slope • San Francisco (CA)

On-site
USD 120,000 - 160,000
Relocation assistance
Flexible hybrid work model
Collaborative work environment