Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
OpenAI is seeking a skilled engineer to build the model runtime within the inference engine that executes frontier models on our custom silicon. You will bridge models on hardware with the cluster serving software, translating workloads into efficient execution while optimizing throughput, latency, and reliability.
You will collaborate across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime, comparable to systems like vLLM but tailored
OpenAI is seeking a skilled engineer to build the model runtime within the inference engine that executes frontier models on our custom silicon. You will bridge models on hardware with the cluster serving software, translating workloads into efficient execution while optimizing throughput, latency, and reliability.
You will collaborate across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime, comparable to systems like vLLM but tailored