Senior Transformer Inference Architect

Oho Group

San Francisco (CA)

On-site

USD 180,000 - 320,000

Full time

24 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Oho Group in San Francisco is seeking a senior inference specialist to architect high-performance transformer inference on a new compute platform. You will optimize model execution, graph transformations and runtime behavior across CPUs, GPUs and accelerators.

You will implement or guide performance-critical C++, CUDA or Triton components, improve attention, GEMM and MoE paths, and design KV-cache, batching and memory management to boost latency, throughput and cost per token.

Qualifications

  • Hands-on optimization experience for transformer or generative-AI inference.
  • Deep experience with low-level implementation in C++, CUDA or Triton.
  • Strong profiling and performance-debugging skills.
  • Ability to optimize across the full inference stack—from models to kernels to runtimes.

Responsibilities

  • Architect high-performance transformer inference on a new compute platform.
  • Optimize model execution, graph transformations and runtime behavior.
  • Develop or guide performance-critical C++, CUDA or Triton components.
  • Improve attention, GEMM, MoE and other critical execution paths.
  • Design KV-cache, batching, memory-management and decoding strategies.
  • Analyze multi-device and multi-node inference performance.
  • Drive improvements in latency, throughput, utilization and cost per token.

Skills

C++
CUDA
Triton
Transformer inference
Profiling
Performance debugging

Tools

CUDA
Triton

Job description

Oho Group in San Francisco is seeking a senior inference specialist to architect high-performance transformer inference on a new compute platform. You will optimize model execution, graph transformations and runtime behavior across CPUs, GPUs and accelerators.

You will implement or guide performance-critical C++, CUDA or Triton components, improve attention, GEMM and MoE paths, and design KV-cache, batching and memory management to boost latency, throughput and cost per token.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Inference Engineer
Distinguished Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000
Edge Transformer Inference Tech Lead
Edge Transformer Inference Tech Lead

OpenAI • San Francisco (CA)

Hybrid
USD 130,000 - 180,000
Inference Software Engineer: High-Performance Transformers
Inference Software Engineer: High-Performance Transformers

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 180,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+2
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Austin (TX)

On-site
USD 191,000 - 239,000
Inference Systems Engineer for Transformers & Low-Latency HPC
Inference Systems Engineer for Transformers & Low-Latency HPC

Etched • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
+1
Senior GPU AI Inference Systems Engineer
Senior GPU AI Inference Systems Engineer

NVIDIA • California (MO)

On-site
USD 196,000 - 288,000
Equity
Comprehensive benefits
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
On-Device AI Inference Engineer — Low-Latency, Embedded
On-Device AI Inference Engineer — Low-Latency, Embedded

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work