Get more replies from employers
Send a job-specific resume in minutes.
OpenAI is seeking a role focused on modeling inference performance across application, model, and fleet layers to drive faster, cheaper deployment. You will translate microbenchmarks into cost-to-serve estimates and build tools for latency, capacity, and cost tradeoffs.
You will collaborate with engineering and research teams to identify bottlenecks, refine performance models, and project how future changes affect inference across production systems.
OpenAI is seeking a role focused on modeling inference performance across application, model, and fleet layers to drive faster, cheaper deployment. You will translate microbenchmarks into cost-to-serve estimates and build tools for latency, capacity, and cost tradeoffs.
You will collaborate with engineering and research teams to identify bottlenecks, refine performance models, and project how future changes affect inference across production systems.