An application made for this job — a tailored resume and cover letter that speak straight to the posting.
TensorX in Dublin is hiring multiple GPU Performance Engineers to optimize inference engines for latency targets. You will patch open-source engines, tune kernels, and bring up new models on day zero within a high-concurrency production environment.
You will work with the Inference Team on CUDA kernels, memory management and KV cache behavior, striving to maximize per-GPU throughput while maintaining correctness and observability.
TensorX in Dublin is hiring multiple GPU Performance Engineers to optimize inference engines for latency targets. You will patch open-source engines, tune kernels, and bring up new models on day zero within a high-concurrency production environment.
You will work with the Inference Team on CUDA kernels, memory management and KV cache behavior, striving to maximize per-GPU throughput while maintaining correctness and observability.