An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Oho Group in the San Francisco area is seeking an experienced inference specialist to optimize the full stack—from transformer models through execution engines, compilers, runtimes and kernels—to multi-accelerator systems for AI workloads.
You will architect high-performance transformer inference on a new compute platform, improve latency and throughput, and contribute hands-on low-level work in C++, CUDA or Triton to push the frontier of AI inference.
We’re supporting an advanced-compute company building a new hardware and software platform for AI workloads.
They’re looking for an inference specialist who can optimize the complete path from transformer models through execution engines, compilers, runtimes and kernels to multi-accelerator systems.