Embedded AI Engineer: On-Device Models & Edge Kernels

Deepgram

San Francisco (CA)

On-site

USD 130,000 - 195,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Deepgram is seeking an Embedded AI Engineer to work on the Partner Platform Engineering team, tackling the lowest layer of the edge stack. You will write and optimize custom kernels and operators for diverse hardware, including embedded SoCs, DSPs, and NPUs, enabling Deepgram models to run on non-NVIDIA accelerators.

Your work will involve quantization, operator fusion, and architecture-specific compilation, with collaboration to fit models to constrained devices and delivery of reusable runtime

Qualifications

  • Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.
  • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
  • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
  • Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
  • A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance.
  • Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.
  • Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.

Responsibilities

  • Write and optimize custom kernels and operators (C, C++, Rust, and platform assembly or intrinsics) for non-NVIDIA accelerators, embedded SoCs, DSPs, and NPUs where the vendor's standard operator set is insufficient for Deepgram models.
  • Own target-side optimization: collapse models onto device execution units through quantization, operator fusion, memory layout, and architecture-specific compilation to meet latency, memory, power, and thermal budgets.
  • Integrate with vendor NPU/DSP toolchains and edge inference runtimes, and extend them with custom operators when the graph doesn't map cleanly.
  • Deliver kernels and runtime components as reusable building blocks that Applied ML Engineers can target when adapting models, with clear interfaces and documented constraints.
  • Build performance-critical runtime code for embedded environments, including embedded Linux, bare-metal, and RTOS targets.
  • Establish per-platform benchmarking and validation for latency, accuracy, power, memory footprint, and utilization, and catch regressions before they ship.
  • Partner with silicon and platform vendors on SDK integration and low-level performance tuning for new chipsets and reference platforms.
  • Feed hardware constraints back to Applied ML and Research so model designs are easier to land on constrained targets.

Skills

C
C++
Rust
Embedded systems
Edge AI
Performance optimization
Hardware-software interaction
RTOS

Tools

ONNX Runtime
TensorRT
TFLite
ExecuTorch

Job description

Deepgram is seeking an Embedded AI Engineer to work on the Partner Platform Engineering team, tackling the lowest layer of the edge stack. You will write and optimize custom kernels and operators for diverse hardware, including embedded SoCs, DSPs, and NPUs, enabling Deepgram models to run on non-NVIDIA accelerators.

Your work will involve quantization, operator fusion, and architecture-specific compilation, with collaboration to fit models to constrained devices and delivery of reusable runtime

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Embedded AI Engineer: Edge Hardware & Kernels
Embedded AI Engineer: Edge Hardware & Kernels

Aimlroles • Northern (KY)

Hybrid
USD 120,000 - 190,000
On-Device AI Engineer: Edge Model Optimization
On-Device AI Engineer: Edge Model Optimization

Deepgram • United States

Remote
USD 140,000 - 210,000
Embedded AI Engineer: On-Device Speech & Edge Inference
Embedded AI Engineer: On-Device Speech & Edge Inference

Deepgram, Inc. • United States

On-site
USD 140,000 - 190,000
Edge ML Engineer: Port Models to Non-NVIDIA & Devices
Edge ML Engineer: Port Models to Non-NVIDIA & Devices

Deepgram • San Francisco (CA)

On-site
USD 150,000 - 230,000
Edge ML Engineer: Port Models to Non-NVIDIA & ARM
Edge ML Engineer: Port Models to Non-NVIDIA & ARM

Deepgram • United States

Remote
USD 150,000 - 230,000
Edge ML Engineer: Port Speech Models to Diverse Hardware
Edge ML Engineer: Port Speech Models to Diverse Hardware

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Embedded AI Engineer, On-Device Models
Embedded AI Engineer, On-Device Models

Deepgram • San Francisco (CA)

On-site
USD 130,000 - 195,000
Edge ML Engineer: Port & Deploy Models to Edge Devices
Edge ML Engineer: Port & Deploy Models to Edge Devices

Precision Labs • Northern (KY)

Hybrid
USD 150,000 - 210,000
Edge ML Engineer: Port & Optimize Models for Edge
Edge ML Engineer: Port & Optimize Models for Edge

Aimlroles • Northern (KY)

Hybrid
USD 150,000 - 210,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000