Director of Engineering, AI Inference Platform

Oho Group

San Francisco (CA)

On-site

USD 260,000 - 420,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Oho Group in San Francisco is seeking a Director of Engineering to build and lead the team responsible for turning its frontier-model inference platform into production-grade software and services.

You will own the architecture and delivery of the end-to-end inference stack, including model enablement, runtimes, scheduling, and serving infrastructure, while mentoring an exceptional team and guiding performance decisions.

Qualifications

  • Significant experience building high-performance inference or distributed AI systems.
  • Leadership experience across inference, ML systems, accelerator software or performance engineering.
  • Strong understanding of transformer inference, attention, KV-cache behaviour, batching and autoregressive decoding.

Responsibilities

  • Define the architecture and roadmap for the end-to-end inference stack.
  • Build and lead teams across inference, runtimes, kernels and distributed systems.
  • Optimise prefill and decode performance across throughput, latency, memory utilisation, power and cost.
  • Develop parallelism and scheduling strategies across chips, servers and racks.
  • Guide work across continuous batching, KV-cache management, speculative decoding and prefill/decode disaggregation.
  • Lead enablement and optimisation of transformer and MoE models.
  • Partner with compiler and kernel teams on graph optimisation, operator fusion, tiling and code generation.
  • Translate production workloads into improvements across compute, memory and interconnect architecture.
  • Establish benchmarking, profiling, observability and performance-regression infrastructure.
  • Work directly with frontier AI companies, hyperscalers and cloud customers.

Skills

High-performance inference
Distributed AI systems
Leadership
Transformer inference
Model parallelism
C++
Python
CUDA

Tools

PyTorch
TensorRT-LLM
JAX
XLA
CUDA
Triton

Job description

Oho Group in San Francisco is seeking a Director of Engineering to build and lead the team responsible for turning its frontier-model inference platform into production-grade software and services.

You will own the architecture and delivery of the end-to-end inference stack, including model enablement, runtimes, scheduling, and serving infrastructure, while mentoring an exceptional team and guiding performance decisions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Manager: Inference Infrastructure Leader
Engineering Manager: Inference Infrastructure Leader

EngineersOfAI • New York (NY), Northern (KY)

Hybrid
USD 230,000 - 360,000
Senior Engineering Manager, AI Inference (LLM Production)
Senior Engineering Manager, AI Inference (LLM Production)

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+13
Senior AI Inference Engineering Lead
Senior AI Inference Engineering Lead

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 230,000 - 300,000
Health benefits
Paid time off
401(k) matching
+1
Head of AI Inference Platform & Strategy
Head of AI Inference Platform & Strategy

FriendliAI • San Francisco (CA)

On-site
USD 210,000 - 280,000
Flexible working hours
Daily meals provided
Unlimited snacks and beverages
+3
Senior AI Infrastructure Engineer - Real-Time Inference
Senior AI Infrastructure Engineer - Real-Time Inference

Didi Labs • San Jose (CA)

Hybrid
USD 170,000 - 351,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Head of AI Inference & Systems Platform
Head of AI Inference & Systems Platform

Blackhornvc • San Francisco (CA)

On-site
USD 220,000 - 420,000
Lead, Inference Infrastructure & Fleet
Lead, Inference Infrastructure & Fleet

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Senior AI Inference Systems Engineer (Equity)
Senior AI Inference Systems Engineer (Equity)

Emploive • San Francisco (CA)

On-site
USD 230,000 - 390,000
Flexible PTO
Medical, dental, vision benefits
Retirement plan
+3