On-Device AI Engineer: Edge Model Optimization

Deepgram

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Deepgram’s embedded AI engineer role focuses on getting speech models running on on-device, low-power hardware across phones, wearables, and edge devices, with real-time, offline-capable capabilities.

You’ll optimize for latency, memory, power, and thermal budgets using quantization, pruning, distillation, and compilation, while collaborating with hardware and research teams to push edge-friendly architectures.

Qualifications

  • Experience delivering production systems on resource-constrained hardware.
  • Strong proficiency in C, C++, and/or Rust.
  • Hands‑on experience with model optimization for on‑device deployment.
  • Familiarity with edge inference runtimes (ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor‑specific NPU/DSP toolchains.
  • A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management.
  • Experience working close to the metal: bare‑metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.
  • Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.

Responsibilities

  • Take Deepgram’s Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
  • Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
  • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
  • Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity.
  • Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
  • Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship.
  • Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
  • Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.

Skills

C/C++/Rust
Communication skills
Builder mindset
Embedded systems

Tools

ONNX Runtime
TensorRT
TFLite
ExecuTorch

Job description

Deepgram’s embedded AI engineer role focuses on getting speech models running on on-device, low-power hardware across phones, wearables, and edge devices, with real-time, offline-capable capabilities.

You’ll optimize for latency, memory, power, and thermal budgets using quantization, pruning, distillation, and compilation, while collaborating with hardware and research teams to push edge-friendly architectures.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Embedded AI Engineer: On-Device Speech & Edge Inference
Embedded AI Engineer: On-Device Speech & Edge Inference

Deepgram, Inc. • United States

On-site
USD 140,000 - 190,000
Embedded AI Engineer: On-Device Models & Edge Kernels
Embedded AI Engineer: On-Device Models & Edge Kernels

Deepgram • San Francisco (CA)

On-site
USD 130,000 - 195,000
Embedded AI Engineer: Edge Hardware & Kernels
Embedded AI Engineer: Edge Hardware & Kernels

Aimlroles • Northern (KY)

Hybrid
USD 120,000 - 190,000
Edge ML Engineer: Port & Optimize Models for Edge
Edge ML Engineer: Port & Optimize Models for Edge

Aimlroles • Northern (KY)

Hybrid
USD 150,000 - 210,000
Edge ML Engineer: Port Models to Non-NVIDIA & Devices
Edge ML Engineer: Port Models to Non-NVIDIA & Devices

Deepgram • San Francisco (CA)

On-site
USD 150,000 - 230,000
Remote Edge AI Engineer — On-Device ML & Optimization
Remote Edge AI Engineer — On-Device ML & Optimization

Bright Vision Technologies • Maple Grove (MN)

Remote
USD 100,000 - 150,000
Edge ML Engineer: Port & Deploy Models to Edge Devices
Edge ML Engineer: Port & Deploy Models to Edge Devices

Precision Labs • Northern (KY)

Hybrid
USD 150,000 - 210,000
Edge ML Engineer: Port Models to Non-NVIDIA & ARM
Edge ML Engineer: Port Models to Non-NVIDIA & ARM

Deepgram • United States

Remote
USD 150,000 - 230,000
Edge ML Engineer: Port Speech Models to Diverse Hardware
Edge ML Engineer: Port Speech Models to Diverse Hardware

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Senior Edge AI Engineer - On-Device ML & Systems
Senior Edge AI Engineer - On-Device ML & Systems

QUALCOMM, Inc. • San Diego (CA)

On-site
USD 141,000 - 211,000
Discretionary bonus
RSU grants
Comprehensive benefits