Head of On-Device AI Inference & Performance

Hark

San Jose (CA)

On-site

USD 300,000 - 500,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hark is an AI company building proactive, multimodal intelligence with advanced hardware and software integration. You will own how our models run on silicon, selecting accelerators, co-designing architectures against latency, memory, and power budgets, and building the low-level inference stack that enables millisecond responsiveness on battery-powered devices.

You will lead the team responsible for implementing efficient transformer execution on new accelerators, collaborating with silicon

Qualifications

  • 8–12+ years in high-performance computing and production workloads on GPUs/NPUs/specialized accelerators.
  • Deep understanding of attention, KV-cache, quantization, and memory bandwidth limits.
  • Designed or optimized inference engines, distributed runtimes, or ML compilers, and wrote kernels when necessary.
  • Experience leading teams on performance-critical software; set direction on a stack, not just contributed.
  • Took a model from research checkpoint to running on constrained hardware in a product.

Responsibilities

  • Evaluate GPUs, NPUs, DSPs, and accelerators for on-device deployment and advise hardware decisions.
  • Collaborate with foundation model and audio ML teams to shape architectures meeting deployment constraints.
  • Build the low-level execution layer, kernels, runtimes, and compiler paths for transformer workloads on target hardware.
  • Partner with silicon vendors and internal hardware teams to bring up accelerators and optimize transformer execution.
  • Hire and lead engineers on performance-critical software and set the bar for the inference stack.

Skills

High-performance computing
Attention & KV-cache
Inference engines / ML compilers
Leadership experience
Model deployment to hardware

Tools

TensorRT
ONNX Runtime
TVM
MLIR

Job description

Hark is an AI company building proactive, multimodal intelligence with advanced hardware and software integration. You will own how our models run on silicon, selecting accelerators, co-designing architectures against latency, memory, and power budgets, and building the low-level inference stack that enables millisecond responsiveness on battery-powered devices.

You will lead the team responsible for implementing efficient transformer execution on new accelerators, collaborating with silicon

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Tech Lead, On-Device AI Inference
Senior Tech Lead, On-Device AI Inference

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
On-Device AI Inference Engineer — Ultra-Low Latency
On-Device AI Inference Engineer — Ultra-Low Latency

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
On-Device AI Inference Engineer — Low-Latency, Embedded
On-Device AI Inference Engineer — Low-Latency, Embedded

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
On-Device AI Inference Engineer San Jose
On-Device AI Inference Engineer San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
On-Device AI Inference Engineer
On-Device AI Inference Engineer

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Technical Lead, On-Device AI Inference
Technical Lead, On-Device AI Inference

Hark • San Jose (CA)

On-site
USD 300,000 - 500,000
Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
Embedded Software Engineer
Embedded Software Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Product Performance Engineer
Product Performance Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Embedded Software Engineer San Jose
Embedded Software Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 120,000 - 300,000