Get more replies from employers
Send a job-specific resume in minutes.
OpenAI's Inference team builds high-volume, low-latency production and research systems around our largest AI models. This role focuses on optimizing performance, resource usage, and reliability in a production environment.
You will partner with researchers and engineers to deploy cutting-edge techniques, improve latency and throughput, and own end-to-end solutions from design to deployment, including tuning on Azure GPUs, distributed stacks, and monitoring tooling.
OpenAI's Inference team builds high-volume, low-latency production and research systems around our largest AI models. This role focuses on optimizing performance, resource usage, and reliability in a production environment.
You will partner with researchers and engineers to deploy cutting-edge techniques, improve latency and throughput, and own end-to-end solutions from design to deployment, including tuning on Azure GPUs, distributed stacks, and monitoring tooling.