Principal Transformer Inference Engineer - High Performance

Oho Group

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

25 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Oho Group in the San Francisco area is seeking an experienced inference specialist to optimize the full stack—from transformer models through execution engines, compilers, runtimes and kernels—to multi-accelerator systems for AI workloads.

You will architect high-performance transformer inference on a new compute platform, improve latency and throughput, and contribute hands-on low-level work in C++, CUDA or Triton to push the frontier of AI inference.

Qualifications

  • Direct optimization of transformer inference required.
  • Hands-on low-level implementation in C++, CUDA or Triton.
  • Experience optimizing across inference stack layers.

Responsibilities

  • Architect high-performance transformer inference on a new compute platform.
  • Optimize model execution, graph transformations and runtime behavior.
  • Develop or guide performance-critical C++, CUDA or Triton components.
  • Improve attention, GEMM, MoE and other critical execution paths.
  • Design KV-cache, batching, memory-management and decoding strategies.
  • Optimize tensor, pipeline and expert parallelism.
  • Analyze multi-device and multi-node inference performance.
  • Drive improvements in latency, throughput, utilization and cost per token.

Skills

Transformer inference optimization
Performance profiling
Low-level programming
System-level optimization

Tools

C++
CUDA
Triton

Job description

Oho Group in the San Francisco area is seeking an experienced inference specialist to optimize the full stack—from transformer models through execution engines, compilers, runtimes and kernels—to multi-accelerator systems for AI workloads.

You will architect high-performance transformer inference on a new compute platform, improve latency and throughput, and contribute hands-on low-level work in C++, CUDA or Triton to push the frontier of AI inference.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Inference Engineer
Principal Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Systems Engineer for Transformers & Low-Latency HPC
Inference Systems Engineer for Transformers & Low-Latency HPC

Etched • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
+1
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Real-Time GPU Optimization Engineer - Inference
Real-Time GPU Optimization Engineer - Inference

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Inference Software Engineer: High-Performance Transformers
Inference Software Engineer: High-Performance Transformers

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 180,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+2
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
GPU Transformer Performance Engineer (Triton/CUDA)
GPU Transformer Performance Engineer (Triton/CUDA)

Luma AI • United States

Remote
USD 180,000 - 280,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000