On-Device AI Inference Engineer San Jose

Hark

San Jose, Northern (CA, KY)

Hybrid

USD 200,000 - 450,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Hark is building multimodal AI hardware and software designed to run transformer workloads with low latency and power constraints. You will write and optimize kernels, manage memory and scheduling across models, and profile real hardware to close the gap between theory and practice.

Ideal candidates have 4-8+ years in performance-critical software for accelerators, strong C/C++, and experience with ONNX Runtime, TVM, MLIR or TensorRT to push inference efficiency on constrained devices.

Qualifications

  • 4-8+ years writing performance-critical software with hands-on optimization on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C/C++ and comfort with SIMD, custom kernels, memory layout, and the profiling tools that go with them.
  • You understand attention, KV-cache behavior, and where transformer inference actually spends its time and memory bandwidth.
  • You reason in compute, memory, and power budgets, and you've optimized against them rather than around them.
  • You've had a model you optimized run in a product on constrained hardware.

Responsibilities

  • Write and optimize the low-level kernels and runtime paths that transformer workloads execute through on target silicon.
  • Decide how multiple models share limited memory and power i.e. residency, scheduling, and swap behavior across concurrent workloads.
  • Profile models on real hardware, find the bottlenecks, and close the gap between theoretical and delivered performance.
  • Take models from full precision to INT8/INT4 and get them running within per-product size, latency, and power budgets.
  • Get transformer workloads executing efficiently on new accelerators as they come online, working alongside the hardware team.
  • Feed real deployment constraints back to the model teams so architecture decisions account for what the hardware can actually do.

Skills

C/C++
Performance optimization
Profiling tools
Memory layout
Compiler/MLIR

Tools

ONNX Runtime
TVM
MLIR
TensorRT

Job description

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

You'll make Hark's models run fast on the hardware we ship. That means writing the kernels, building the runtime paths, and profiling transformer workloads on DSPs, NPUs, and other constrained targets until they hit the latency and power budgets our devices are built around. This is hands-on systems work close to the metal, on a small team where the code you write is what users feel as response time.

Responsibilities

  • Write and optimize the low-level kernels and runtime paths that transformer workloads execute through on target silicon.
  • Decide how multiple models share limited memory and power i.e. residency, scheduling, and swap behavior across concurrent workloads.
  • Profile models on real hardware, find the bottlenecks, and close the gap between theoretical and delivered performance.
  • Take models from full precision to INT8/INT4 and get them running within per-product size, latency, and power budgets.
  • Get transformer workloads executing efficiently on new accelerators as they come online, working alongside the hardware team.
  • Feed real deployment constraints back to the model teams so architecture decisions account for what the hardware can actually do.

Requirements

  • 4-8+ years writing performance-critical software, with hands-on optimization on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C/C++ and comfort with SIMD, custom kernels, memory layout, and the profiling tools that go with them.
  • You understand attention, KV-cache behavior, and where transformer inference actually spends its time and memory bandwidth.
  • You reason in compute, memory, and power budgets, and you've optimized against them rather than around them.
  • You've had a model you optimized run in a product on constrained hardware.

Bonus Qualifications

  • Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
  • Background in speech or audio inference, where latency is perceptible to the user.

Compensation

The US base salary range for this full-time position is between $200,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Device AI Inference Engineer
On-Device AI Inference Engineer

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
Technical Lead, On-Device AI Inference
Technical Lead, On-Device AI Inference

Hark • San Jose (CA)

On-site
USD 300,000 - 500,000
Data Engineering Lead San Jose
Data Engineering Lead San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 170,000 - 450,000
Embedded Software Engineer
Embedded Software Engineer

Hark • San Jose (CA)

On-site
USD 120,000 - 300,000
Embedded Software Engineer San Jose
Embedded Software Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 120,000 - 300,000
Product Performance Engineer San Jose
Product Performance Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 120,000 - 300,000
Full-Stack Engineer San Jose
Full-Stack Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 170,000 - 400,000
Member of Technical Staff, Multimodal San Jose
Member of Technical Staff, Multimodal San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Mobile Android Engineer San Jose
Mobile Android Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000