Turn this role into an interview — a resume and cover letter built around what this employer wants.
OpenAI seeks engineers to design and operate Kubernetes-based controllers that coordinate GPU compute infrastructure across sites. You will define APIs, build provisioning services, and manage lifecycle operations from boot to decommissioning, ensuring reliability and scalability.
Ideal candidates have strong distributed-systems background, experience with Kubernetes reconciliation, and the ability to troubleshoot complex cross-boundary issues with staged rollouts across nodes and racks.
OpenAI's Compute Foundations team builds software that manages GPU compute infrastructure across data centers and sites, supporting model training and inference. In this role, you will design and operate Kubernetes-based distributed systems that provision, configure, and manage compute resources throughout their lifecycle, connecting global services with bare-metal systems management.