Stand out for this role — generate a tailored resume and cover letter in about a minute.
Inferact Singapore is seeking a TPU performance engineer to advance vLLM as a premier inference engine on Google TPUs. You will develop and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling to achieve frontier inference performance on TPU hardware.
You will work at the boundary of inference systems, kernels, compilers, and hardware architecture, delivering production-ready model serving with clear correctness,
Inferact Singapore is seeking a TPU performance engineer to advance vLLM as a premier inference engine on Google TPUs. You will develop and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling to achieve frontier inference performance on TPU hardware.
You will work at the boundary of inference systems, kernels, compilers, and hardware architecture, delivering production-ready model serving with clear correctness,