On-Device AI Inference Engineer — Low-Latency, Embedded

Hark

San Jose (CA)

On-site

USD 200,000 - 450,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hark is seeking an experienced systems engineer to optimize transformer workloads on dedicated hardware, writing kernels and runtime paths for latency and power targets. You will push models from full precision toward INT8/INT4 while coordinating with hardware teams to maximize efficiency.

The role requires 4–8+ years in performance-critical software, strong C/C++, SIMD, and deep knowledge of attention/KV-cache dynamics within deployment constraints.

Qualifications

  • 4-8+ years writing performance-critical software and hands-on optimization.
  • Strong C/C++ and familiarity with SIMD, custom kernels, memory layout.
  • Understanding of attention, KV-cache behavior, and transformer inference timing.
  • Experience optimizing software for constrained hardware (GPUs/NPUs/DSPs).

Responsibilities

  • Write and optimize low-level kernels and runtime paths for transformer workloads on target silicon.
  • Manage memory residency, scheduling, and swap behavior across concurrent models.
  • Profile models on real hardware and eliminate bottlenecks to close theory-to-performance gaps.
  • Convert full-precision models to INT8/INT4 within size, latency, and power budgets.
  • Collaborate with hardware teams to enable new accelerators and feed deployment constraints back to model teams.
  • Ensure transformer workloads run efficiently on hardware accelerators as they come online.

Skills

C/C++
SIMD
Performance optimization
Profiling
Transformer inference

Tools

ONNX Runtime
TVM
MLIR
TensorRT

Job description

Hark is seeking an experienced systems engineer to optimize transformer workloads on dedicated hardware, writing kernels and runtime paths for latency and power targets. You will push models from full precision toward INT8/INT4 while coordinating with hardware teams to maximize efficiency.

The role requires 4–8+ years in performance-critical software, strong C/C++, SIMD, and deep knowledge of attention/KV-cache dynamics within deployment constraints.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Device AI Inference Engineer — Ultra-Low Latency
On-Device AI Inference Engineer — Ultra-Low Latency

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
Head of On-Device AI Inference & Performance
Head of On-Device AI Inference & Performance

Hark • San Jose (CA)

On-site
USD 300,000 - 500,000
Senior Tech Lead, On-Device AI Inference
Senior Tech Lead, On-Device AI Inference

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
On-Device AI Inference Engineer San Jose
On-Device AI Inference Engineer San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
On-Device AI Inference Engineer
On-Device AI Inference Engineer

Hark • San Jose (CA)

On-site
USD 200,000 - 450,000
Inference Systems Engineer for Transformers & Low-Latency HPC
Inference Systems Engineer for Transformers & Low-Latency HPC

Etched • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
+1
Distinguished Inference Engineer
Distinguished Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000
Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
Technical Lead, On-Device AI Inference
Technical Lead, On-Device AI Inference

Hark • San Jose (CA)

On-site
USD 300,000 - 500,000
AI Inference Engineer – High-Performance GPU Systems
AI Inference Engineer – High-Performance GPU Systems

Perplexity • California (MO)

On-site
USD 120,000 - 170,000