An application made for this job — a tailored resume and cover letter that speak straight to the posting.
techire ai is seeking a GPU Optimisation Engineer for real-time inference in production AI workloads. The role focuses on pushing GPU performance to sub-50ms latency under high concurrency, close to the metal across kernel and runtime layers.
You will profile and optimize large generative models, write custom CUDA/Triton kernels, and collaborate with research to deliver production-ready inference at scale. SF relocation/visa sponsorship available.
techire ai is seeking a GPU Optimisation Engineer for real-time inference in production AI workloads. The role focuses on pushing GPU performance to sub-50ms latency under high concurrency, close to the metal across kernel and runtime layers.
You will profile and optimize large generative models, write custom CUDA/Triton kernels, and collaborate with research to deliver production-ready inference at scale. SF relocation/visa sponsorship available.