Distinguished Inference Engineer

Oho Group

San Francisco (CA)

On-site

USD 180,000 - 320,000

Full time

17 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Oho Group in San Francisco is seeking a senior inference specialist to architect high-performance transformer inference on a new compute platform. You will optimize model execution, graph transformations and runtime behavior across CPUs, GPUs and accelerators.

You will implement or guide performance-critical C++, CUDA or Triton components, improve attention, GEMM and MoE paths, and design KV-cache, batching and memory management to boost latency, throughput and cost per token.

Qualifications

  • Hands-on optimization experience for transformer or generative-AI inference.
  • Deep experience with low-level implementation in C++, CUDA or Triton.
  • Strong profiling and performance-debugging skills.
  • Ability to optimize across the full inference stack—from models to kernels to runtimes.

Responsibilities

  • Architect high-performance transformer inference on a new compute platform.
  • Optimize model execution, graph transformations and runtime behavior.
  • Develop or guide performance-critical C++, CUDA or Triton components.
  • Improve attention, GEMM, MoE and other critical execution paths.
  • Design KV-cache, batching, memory-management and decoding strategies.
  • Analyze multi-device and multi-node inference performance.
  • Drive improvements in latency, throughput, utilization and cost per token.

Skills

C++
CUDA
Triton
Transformer inference
Profiling
Performance debugging

Tools

CUDA
Triton

Job description

We’re supporting an advanced-compute company building a new hardware and software platform for AI workloads.

They’re looking for an inference specialist who can optimize the complete path from transformer models through execution engines, compilers, runtimes and kernels to multi-accelerator systems.

What you’ll work on
  • Architect high-performance transformer inference on a new compute platform
  • Optimize model execution, graph transformations and runtime behavior
  • Develop or guide performance-critical C++, CUDA or Triton components
  • Improve attention, GEMM, MoE and other critical execution paths
  • Design KV-cache, batching, memory-management and decoding strategies
  • Optimize tensor, pipeline and expert parallelism
  • Analyze multi-device and multi-node inference performance
  • Drive improvements in latency, throughput, utilization and cost per token
What we’re looking for
  • Principal, Distinguished or equivalent senior technical scope
  • Direct optimization of transformer or generative-AI inference
  • Hands-on low-level implementation in C++, CUDA, Triton or similar
  • Deep expertise across multiple connected layers of the inference stack
  • Strong profiling and performance-debugging skills
  • Experience taking optimizations beyond isolated kernels into complete systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Transformer Inference Architect
Senior Transformer Inference Architect

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Engineering Lead, Inference Optimization
Engineering Lead, Inference Optimization

Shields Group Search • United States

On-site
USD 270,000 - 330,000
Equity
Crypto token compensation
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • California (MO)

On-site
USD 120,000 - 170,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2