Software Engineer, Workload Enablement

OpenAI

United States

Remote

USD 140,000 - 220,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

OpenAI is seeking an SW Engineer to enable production workloads and end-to-end testing on new platforms. You will port inference and training workloads to early-access systems, analyze performance bottlenecks, and characterize end-to-end behavior of compute, comms, storage and control planes.

You will develop test harnesses, CI-ready benchmarks, and ensure scalability with containerization, Kubernetes integration, and telemetry hooks, while collaborating with vendors and internal teams to drive

Qualifications

  • Experience porting inference and training workloads to new platforms.

Responsibilities

  • Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar.
  • Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts.
  • Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects
  • Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection).
  • Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops).
  • Work cross-functionally with vendors and internal stakeholders by producing

Job description

About the Team

The Scaling team is responsible for the architectural and engineering backbone of OpenAI's infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization.

About the Role

We're hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes).

Key Responsibilities
  • Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar.
  • Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts.
  • Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects
  • Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection).
  • Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops).
  • Work cross-functionally with vendors and internal stakeholders by producing
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Workload Enablement
Software Engineer, Workload Enablement

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Equal Opportunity Employer
Commitment to diversity and inclusion
Reasonable accommodations available
Software Engineer, Workload Enablement
Software Engineer, Workload Enablement

OpenAI • Seattle (WA)

On-site
USD 293,000 - 455,000
Production Workloads Engineer: Benchmark Enablement
Production Workloads Engineer: Benchmark Enablement

OpenAI • United States

Remote
USD 140,000 - 220,000
Workload Porting & Performance Engineer (AI/ML)
Workload Porting & Performance Engineer (AI/ML)

OpenAI • Seattle (WA)

Hybrid
USD 170,000 - 260,000
Relocation assistance
Hybrid work model
Workload Porting & Performance Engineer
Workload Porting & Performance Engineer

OpenAI • Seattle (WA)

On-site
USD 170,000 - 260,000
Relocation assistance
Hybrid work model
Platform Workload Engineer: Performance & E2E Testing
Platform Workload Engineer: Performance & E2E Testing

OpenAI • Seattle (WA)

On-site
USD 293,000 - 455,000
Workload Porting & Performance Engineer
Workload Porting & Performance Engineer

OpenAI • Los Angeles (CA)

On-site
USD 342,000 - 555,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

OpenAI • United States

Remote
USD 180,000 - 240,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

OpenAI • Seattle (WA)

On-site
USD 180,000 - 240,000