Technical Lead, On-Device AI Inference

Hark

San Jose (CA)

On-site

USD 300,000 - 500,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Hark is an AI company building proactive, multimodal intelligence with advanced hardware and software integration. You will own how our models run on silicon, selecting accelerators, co-designing architectures against latency, memory, and power budgets, and building the low-level inference stack that enables millisecond responsiveness on battery-powered devices.

You will lead the team responsible for implementing efficient transformer execution on new accelerators, collaborating with silicon

Qualifications

  • 8–12+ years in high-performance computing and production workloads on GPUs/NPUs/specialized accelerators.
  • Deep understanding of attention, KV-cache, quantization, and memory bandwidth limits.
  • Designed or optimized inference engines, distributed runtimes, or ML compilers, and wrote kernels when necessary.
  • Experience leading teams on performance-critical software; set direction on a stack, not just contributed.
  • Took a model from research checkpoint to running on constrained hardware in a product.

Responsibilities

  • Evaluate GPUs, NPUs, DSPs, and accelerators for on-device deployment and advise hardware decisions.
  • Collaborate with foundation model and audio ML teams to shape architectures meeting deployment constraints.
  • Build the low-level execution layer, kernels, runtimes, and compiler paths for transformer workloads on target hardware.
  • Partner with silicon vendors and internal hardware teams to bring up accelerators and optimize transformer execution.
  • Hire and lead engineers on performance-critical software and set the bar for the inference stack.

Skills

High-performance computing
Attention & KV-cache
Inference engines / ML compilers
Leadership experience
Model deployment to hardware

Tools

TensorRT
ONNX Runtime
TVM
MLIR

Job description

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We’re pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today’s AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we’re developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

You’ll own how Hark’s models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You’ll build and lead the team that does it. The ceiling on what our hardware can do is set here.

Responsibilities
  • Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device deployment, and own the recommendation hardware decisions are made against.
  • Work with the foundation model and audio ML teams to shape architectures that meet deployment constraints before training locks them in.
  • Build the low-level execution layer, custom kernels, runtime systems, and compiler paths that transformer workloads run through on target hardware.
  • Partner with silicon vendors and internal hardware teams to bring up new accelerators and get efficient transformer execution on them early.
  • Hire and lead a team of engineers on performance-critical software, and set the technical bar for the inference stack.
Requirements
  • 8–12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators.
  • Deep understanding of attention, KV-cache behavior, quantization effects, and memory bandwidth limits.
  • You’ve designed or optimized inference engines, distributed runtimes, or ML compilers, and you write the kernels yourself when it matters.
  • Experience leading teams on performance-critical software. You’ve set direction on a stack, not just contributed to one.
  • You’ve taken a model from a research checkpoint to running on constrained hardware in a product people use.
Bonus Qualifications
  • Hands-on experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Experience with speech, audio, or streaming multimodal inference where latency is perceptible to the user.
  • Contributions to open-source inference or compiler toolchains (TensorRT, ONNX Runtime, TVM, MLIR, and similar).
Compensation

The US base salary range for this full-time position is between $300,000 - $500,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
On-Device AI Inference Engineer
On-Device AI Inference Engineer

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
On-Device AI Inference Engineer San Jose
On-Device AI Inference Engineer San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
Product Performance Engineer
Product Performance Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Embedded Software Engineer
Embedded Software Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Data Engineering Lead
Data Engineering Lead

Hark • San Jose (CA)

On-site
USD 170,000 - 450,000
Lead Audio ML Engineer
Lead Audio ML Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Platform Engineer
Platform Engineer

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000
Infrastructure, Large-scale Training
Infrastructure, Large-scale Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Full-Stack Engineer
Full-Stack Engineer

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000