On-Device AI Inference Engineer — Ultra-Low Latency

Hark

San Jose, Northern (CA, KY)

Hybrid

USD 200,000 - 450,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Hark is building multimodal AI hardware and software designed to run transformer workloads with low latency and power constraints. You will write and optimize kernels, manage memory and scheduling across models, and profile real hardware to close the gap between theory and practice.

Ideal candidates have 4-8+ years in performance-critical software for accelerators, strong C/C++, and experience with ONNX Runtime, TVM, MLIR or TensorRT to push inference efficiency on constrained devices.

Qualifications

  • 4-8+ years writing performance-critical software with hands-on optimization on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C/C++ and comfort with SIMD, custom kernels, memory layout, and the profiling tools that go with them.
  • You understand attention, KV-cache behavior, and where transformer inference actually spends its time and memory bandwidth.
  • You reason in compute, memory, and power budgets, and you've optimized against them rather than around them.
  • You've had a model you optimized run in a product on constrained hardware.

Responsibilities

  • Write and optimize the low-level kernels and runtime paths that transformer workloads execute through on target silicon.
  • Decide how multiple models share limited memory and power i.e. residency, scheduling, and swap behavior across concurrent workloads.
  • Profile models on real hardware, find the bottlenecks, and close the gap between theoretical and delivered performance.
  • Take models from full precision to INT8/INT4 and get them running within per-product size, latency, and power budgets.
  • Get transformer workloads executing efficiently on new accelerators as they come online, working alongside the hardware team.
  • Feed real deployment constraints back to the model teams so architecture decisions account for what the hardware can actually do.

Skills

C/C++
Performance optimization
Profiling tools
Memory layout
Compiler/MLIR

Tools

ONNX Runtime
TVM
MLIR
TensorRT

Job description

Hark is building multimodal AI hardware and software designed to run transformer workloads with low latency and power constraints. You will write and optimize kernels, manage memory and scheduling across models, and profile real hardware to close the gap between theory and practice.

Ideal candidates have 4-8+ years in performance-critical software for accelerators, strong C/C++, and experience with ONNX Runtime, TVM, MLIR or TensorRT to push inference efficiency on constrained devices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Device AI Inference Engineer — Low-Latency, Embedded
On-Device AI Inference Engineer — Low-Latency, Embedded

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Senior Tech Lead, On-Device AI Inference
Senior Tech Lead, On-Device AI Inference

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
On-Device AI Inference Engineer San Jose
On-Device AI Inference Engineer San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
On-Device AI Inference Engineer
On-Device AI Inference Engineer

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Distinguished Inference Engineer
Distinguished Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000