An application made for this job — a tailored resume and cover letter that speak straight to the posting.
OpenAI is seeking an SW Engineer to enable production workloads and end-to-end testing on new platforms. You will port inference and training workloads to early-access systems, analyze performance bottlenecks, and characterize end-to-end behavior of compute, comms, storage and control planes.
You will develop test harnesses, CI-ready benchmarks, and ensure scalability with containerization, Kubernetes integration, and telemetry hooks, while collaborating with vendors and internal teams to drive
The Scaling team is responsible for the architectural and engineering backbone of OpenAI's infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization.
We're hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes).