Director of Engineering

Oho Group

San Francisco (CA)

On-site

USD 260,000 - 420,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oho Group in San Francisco is seeking a Director of Engineering to build and lead the team responsible for turning its frontier-model inference platform into production-grade software and services.

You will own the architecture and delivery of the end-to-end inference stack, including model enablement, runtimes, scheduling, and serving infrastructure, while mentoring an exceptional team and guiding performance decisions.

Qualifications

  • Significant experience building high-performance inference or distributed AI systems.
  • Leadership experience across inference, ML systems, accelerator software or performance engineering.
  • Strong understanding of transformer inference, attention, KV-cache behaviour, batching and autoregressive decoding.

Responsibilities

  • Define the architecture and roadmap for the end-to-end inference stack.
  • Build and lead teams across inference, runtimes, kernels and distributed systems.
  • Optimise prefill and decode performance across throughput, latency, memory utilisation, power and cost.
  • Develop parallelism and scheduling strategies across chips, servers and racks.
  • Guide work across continuous batching, KV-cache management, speculative decoding and prefill/decode disaggregation.
  • Lead enablement and optimisation of transformer and MoE models.
  • Partner with compiler and kernel teams on graph optimisation, operator fusion, tiling and code generation.
  • Translate production workloads into improvements across compute, memory and interconnect architecture.
  • Establish benchmarking, profiling, observability and performance-regression infrastructure.
  • Work directly with frontier AI companies, hyperscalers and cloud customers.

Skills

High-performance inference
Distributed AI systems
Leadership
Transformer inference
Model parallelism
C++
Python
CUDA

Tools

PyTorch
TensorRT-LLM
JAX
XLA
CUDA
Triton

Job description

A heavily funded AI hardware company is building a vertically integrated platform for frontier-model inference, spanning custom silicon, rack-scale systems, interconnects, compilers, runtimes and serving software.

Its first production silicon has returned and the team is now validating rack-scale systems with major AI customers. The platform targets demanding workloads including large mixture-of-experts models, long-context inference and agentic applications, with the goal of substantially improving throughput, latency, power efficiency and cost per token.

The Role

The company is hiring a Director of Engineering to build and lead the organisation responsible for turning its custom accelerator into a production-grade inference platform.

You will own the technical direction and delivery of the inference stack across model enablement, distributed execution, runtime scheduling, kernel performance and serving infrastructure. This is a hands-on leadership position requiring someone capable of building an exceptional team while remaining closely involved in architecture and performance decisions.

Responsibilities
  • Define the architecture and roadmap for the end-to-end inference stack.
  • Build and lead teams across inference, runtimes, kernels and distributed systems.
  • Optimise prefill and decode performance across throughput, latency, memory utilisation, power and cost.
  • Develop parallelism and scheduling strategies across chips, servers and racks.
  • Guide work across continuous batching, KV-cache management, speculative decoding and prefill/decode disaggregation.
  • Lead enablement and optimisation of transformer and MoE models.
  • Partner with compiler and kernel teams on graph optimisation, operator fusion, tiling and code generation.
  • Translate production workloads into improvements across compute, memory and interconnect architecture.
  • Establish benchmarking, profiling, observability and performance-regression infrastructure.
  • Work directly with frontier AI companies, hyperscalers and cloud customers.
Requirements
  • Significant experience building high-performance inference or distributed AI systems.
  • Leadership experience across inference, ML systems, accelerator software or performance engineering.
  • Strong understanding of transformer inference, attention, KV-cache behaviour, batching and autoregressive decoding.
  • Experience scaling large models across multiple GPUs or AI accelerators.
  • Knowledge of tensor, pipeline and expert parallelism.
  • Experience with vLLM, SGLang, TensorRT-LLM, PyTorch, JAX, XLA or similar systems.
  • Understanding of accelerator architecture, memory hierarchies, interconnects and compute utilisation.
  • Technical strength in C++, Python and CUDA, Triton, ROCm or another low-level accelerator environment.
  • Ability to diagnose performance across frameworks, compilers, runtimes, kernels and hardware.
  • A track record of recruiting and leading highly capable engineering teams.
Particularly Relevant Experience
  • Leading inference or ML-systems teams at a frontier AI lab, hyperscaler, GPU company or accelerator startup.
  • Bringing up models on new silicon before the surrounding software ecosystem was mature.
  • Developing kernels for attention, GEMM, MoE, communication or quantised inference.
  • Building production systems for low-latency or high-throughput LLM serving.

This is an opportunity to define the inference organisation around a new computing platform rather than inherit an established stack. You will influence both the software and silicon roadmap while helping move a technically ambitious architecture into large-scale customer deployment.

A highly competitive compensation and equity package is available.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Tech Lead Manager, Inference
Tech Lead Manager, Inference

lumalabs-ai • San Francisco (CA)

On-site
USD 230,000 - 350,000
Principal Inference Engineer
Principal Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 240,000
Applied Researcher – AI Expert
Applied Researcher – AI Expert

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 280,000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 280,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Applied Researcher – AI Expert
Applied Researcher – AI Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Member of Technical Staff (Inference) - AI Infrastructure
Member of Technical Staff (Inference) - AI Infrastructure

Hamilton Barnes • United States

On-site
USD 225,000 - 275,000
Full Benefits
Staff+ Software Engineer, Inference Velocity
Staff+ Software Engineer, Inference Velocity

Anthropic • San Francisco (CA)

On-site
USD 405,000 - 485,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave