Turn this role into an interview — a resume and cover letter built around what this employer wants.
techire ai is seeking a GPU Optimisation Engineer for real-time inference in production AI workloads. The role focuses on pushing GPU performance to sub-50ms latency under high concurrency, close to the metal across kernel and runtime layers.
You will profile and optimize large generative models, write custom CUDA/Triton kernels, and collaborate with research to deliver production-ready inference at scale. SF relocation/visa sponsorship available.
Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal workloads?
This team is building low-latency AI systems where milliseconds actually matter. The target isn’t “faster than baseline.” It’s sub-50ms time-to-first-token at 100+ concurrent requests on a single H100 — while maintaining model quality.
They’re hiring a GPU Optimisation Engineer who understands GPUs at an architectural level. Someone who knows where performance is really lost: memory hierarchy, kernel launch overhead, occupancy limits, scheduling inefficiencies, KV cache behaviour, attention paths. The work sits close to the metal, inside inference execution — not general infra, not model research.
You’ll operate across the kernel and runtime layers, profiling large-scale speech and multimodal models end-to‑end and removing bottlenecks wherever they appear.
This is hands‑on optimisation work across the stack. No layers of bureaucracy. No “platform ownership” theatre. Just deep performance engineering applied to models that are actively evolving.
The company is revenue‑generating, its models are used by global enterprises, and the SF R&D team is expanding following a recent raise. This is growth hiring, not backfill.
If you care about real-time constraints, GPU architecture, and squeezing every last millisecond out of large models, this is worth a conversation.
All applicants will receive a response.