An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Oho Group in San Francisco is seeking a senior inference specialist to architect high-performance transformer inference on a new compute platform. You will optimize model execution, graph transformations and runtime behavior across CPUs, GPUs and accelerators.
You will implement or guide performance-critical C++, CUDA or Triton components, improve attention, GEMM and MoE paths, and design KV-cache, batching and memory management to boost latency, throughput and cost per token.
We’re supporting an advanced-compute company building a new hardware and software platform for AI workloads.
They’re looking for an inference specialist who can optimize the complete path from transformer models through execution engines, compilers, runtimes and kernels to multi-accelerator systems.