Senior Tech Lead, On-Device AI Inference

Hark

San Jose, Northern (CA, KY)

Hybrid

USD 300,000 - 500,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hark is seeking a senior hardware/software leader to own how our models run on silicon. You will select accelerators, shape architectures within latency, memory, and power budgets, and build the low-level inference stack that enables millisecond responses on battery-powered devices.

You will lead a team of engineers, partner with silicon vendors, and drive efficient transformer execution on new accelerators from research to product.

Qualifications

  • 8–12+ years in high-performance computing with production workloads on GPUs or accelerators.
  • Deep understanding of attention, KV-cache, quantization, and memory bandwidth limits.
  • You've designed or optimized inference engines, runtimes, or ML compilers and write kernels when needed.
  • You've led teams on performance-critical software and set architectural direction.

Responsibilities

  • Evaluate GPUs, NPUs, DSPs, and accelerators for on-device deployment and advocate for hardware decisions.
  • Collaborate with foundation model and audio ML teams to meet deployment constraints before training.
  • Build low-level execution layer, kernels, runtimes, and compiler paths for transformer workloads on target hardware.
  • Partner with silicon vendors and internal teams to bring up accelerators and optimize transformer execution.
  • Hire and lead engineers, setting the technical bar for the inference stack.

Skills

High-performance computing
Team leadership
ML inference
GPU/accelerator knowledge
Performance optimization

Tools

TensorRT
ONNX Runtime
TVM
MLIR

Job description

Hark is seeking a senior hardware/software leader to own how our models run on silicon. You will select accelerators, shape architectures within latency, memory, and power budgets, and build the low-level inference stack that enables millisecond responses on battery-powered devices.

You will lead a team of engineers, partner with silicon vendors, and drive efficient transformer execution on new accelerators from research to product.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of On-Device AI Inference & Performance
Head of On-Device AI Inference & Performance

Hark • San Jose (CA)

On-site
USD 300,000 - 500,000
On-Device AI Inference Engineer — Ultra-Low Latency
On-Device AI Inference Engineer — Ultra-Low Latency

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
On-Device AI Inference Engineer — Low-Latency, Embedded
On-Device AI Inference Engineer — Low-Latency, Embedded

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
Technical Lead, On-Device AI Inference
Technical Lead, On-Device AI Inference

Hark • San Jose (CA)

On-site
USD 300,000 - 500,000
On-Device AI Inference Engineer San Jose
On-Device AI Inference Engineer San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
On-Device AI Inference Engineer
On-Device AI Inference Engineer

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Senior ASIC Design Engineer — AI Inference Hardware
Senior ASIC Design Engineer — AI Inference Hardware

Tensordyne • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
Meals
Snacks & drinks
Unlimited PTO
+1
Lead, Silicon & Systems for AI Inference Hardware
Lead, Silicon & Systems for AI Inference Hardware

Positron • United States

Remote
USD 200,000 - 350,000
Health insurance
Life & disability coverage
Remote-first culture
+1
On-Device Transformer Tech Lead — Edge AI Inference
On-Device Transformer Tech Lead — Edge AI Inference

OpenAI • California (MO)

Hybrid
USD 230,000 - 320,000
Relocation assistance
Hybrid work model