Stand out for this role — generate a tailored resume and cover letter in about a minute.
TensorX, based in Dublin, is seeking GPU Performance Engineers to optimize inference engines, patch open-source components, and push improvements across a live fleet. You will work on kernel-level performance, model bring-up, and cache behavior to meet tight latency targets in a high-concurrency environment.
Reporting to the CTO, you’ll patch defects in SGLang and related tools, verify fixes on production traffic, and collaborate with the Inference Team to size pools and improve throughput.
TensorX, based in Dublin, is seeking GPU Performance Engineers to optimize inference engines, patch open-source components, and push improvements across a live fleet. You will work on kernel-level performance, model bring-up, and cache behavior to meet tight latency targets in a high-concurrency environment.
Reporting to the CTO, you’ll patch defects in SGLang and related tools, verify fixes on production traffic, and collaborate with the Inference Team to size pools and improve throughput.