A complete application in a minute — tailored resume and cover letter, ready to send.
Oho Group in San Francisco is seeking a senior inference specialist to architect high-performance transformer inference on a new compute platform. You will optimize model execution, graph transformations and runtime behavior across CPUs, GPUs and accelerators.
You will implement or guide performance-critical C++, CUDA or Triton components, improve attention, GEMM and MoE paths, and design KV-cache, batching and memory management to boost latency, throughput and cost per token.
Oho Group in San Francisco is seeking a senior inference specialist to architect high-performance transformer inference on a new compute platform. You will optimize model execution, graph transformations and runtime behavior across CPUs, GPUs and accelerators.
You will implement or guide performance-critical C++, CUDA or Triton components, improve attention, GEMM and MoE paths, and design KV-cache, batching and memory management to boost latency, throughput and cost per token.