Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Oho Group in the San Francisco area is seeking an experienced inference specialist to optimize the full stack—from transformer models through execution engines, compilers, runtimes and kernels—to multi-accelerator systems for AI workloads.
You will architect high-performance transformer inference on a new compute platform, improve latency and throughput, and contribute hands-on low-level work in C++, CUDA or Triton to push the frontier of AI inference.
Oho Group in the San Francisco area is seeking an experienced inference specialist to optimize the full stack—from transformer models through execution engines, compilers, runtimes and kernels—to multi-accelerator systems for AI workloads.
You will architect high-performance transformer inference on a new compute platform, improve latency and throughput, and contribute hands-on low-level work in C++, CUDA or Triton to push the frontier of AI inference.