Edge ML Engineer: Port Models to Non-NVIDIA & Devices

Deepgram

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Deepgram is seeking an Applied ML Engineer for the Partner Platform Engineering team to port and optimize speech models for edge devices. You’ll adapt models to run on non-NVIDIA hardware, adjust kernels, and validate accuracy and latency on real devices.

You will collaborate with Embedded AI Engineers, run benchmarks, and own the deployment path from model to edge target, enabling repeatable shipping across platforms.

Qualifications

  • Hands-on experience deploying ML models to edge or non-NVIDIA hardware in production.
  • Working knowledge of quantization and precision tradeoffs (INT8, FP16, mixed precision, calibration).
  • Experience with at least one edge or vendor inference runtime and its conversion toolchain (ONNX Runtime, TFLite, ExecuTorch, OpenVINO, Qualcomm AI Engine, or a vendor NPU SDK).
  • Ability to modify a model to fit a platform: reading and rewriting model graphs, swapping unsupported operators, and adjusting architecture parameters without breaking accuracy.
  • Strong Python and PyTorch, production-quality engineering habits: tests, reproducibility, and benchmarks that others can rerun.
  • Comfort building automation around model conversion and deployment.
  • A builder mindset and clear communication: you can scope a port on an unfamiliar platform and drive it to a measured result.

Responsibilities

  • Port Deepgram speech models to non-NVIDIA and edge platforms, adapting model structure and parameters so they run within the target's existing operator set, runtime, and kernels with minimal modification.
  • Own serving-side model decisions for edge targets: quantization and precision choices, operator substitution, graph rewrites, and architecture tweaks that fit a model to a device's constraints while holding accuracy and latency.
  • Validate every port on real hardware: build accuracy, latency, throughput, and memory benchmarks per platform, and catch regressions before a customer does.
  • Build the deployment path for edge targets: model packaging, conversion pipelines, versioning, and automated delivery so shipping a model to a new device is repeatable rather than bespoke.
  • Work with Embedded AI Engineers when a standard kernel isn't enough: specify what the model needs, then adapt the model to use the custom kernel they deliver.
  • Partner with platform and silicon vendors on their runtimes and toolchains, and turn their expected model format and operator conventions into a working Deepgram deployment.
  • Feed edge constraints back to Research and Impeller so future models are easier to port, without taking on research or core productionization work yourself.
  • As the team grows, take on adjacent production concerns at the edge: automated deployment, model security and integrity on customer hardware, and fleet-level observability.

Skills

Edge deployment
Python
PyTorch
Quantization
Inference runtimes
Model optimization
Automation tooling

Tools

ONNX Runtime
TFLite
OpenVINO
ExecuTorch
Vendor NPU SDK

Job description

Deepgram is seeking an Applied ML Engineer for the Partner Platform Engineering team to port and optimize speech models for edge devices. You’ll adapt models to run on non-NVIDIA hardware, adjust kernels, and validate accuracy and latency on real devices.

You will collaborate with Embedded AI Engineers, run benchmarks, and own the deployment path from model to edge target, enabling repeatable shipping across platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Edge ML Engineer: Port & Optimize Models for Edge
Edge ML Engineer: Port & Optimize Models for Edge

Aimlroles • Northern (KY)

Hybrid
USD 150,000 - 210,000
Edge ML Engineer: Port Models to Non-NVIDIA & ARM
Edge ML Engineer: Port Models to Non-NVIDIA & ARM

Deepgram • United States

Remote
USD 150,000 - 230,000
Edge ML Engineer: Port Speech Models to Diverse Hardware
Edge ML Engineer: Port Speech Models to Diverse Hardware

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Edge ML Engineer: Port & Deploy Models to Edge Devices
Edge ML Engineer: Port & Deploy Models to Edge Devices

Precision Labs • Northern (KY)

Hybrid
USD 150,000 - 210,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Precision Labs • Northern (KY)

Hybrid
USD 150,000 - 210,000
Embedded AI Engineer: On-Device Models & Edge Kernels
Embedded AI Engineer: On-Device Models & Edge Kernels

Deepgram • San Francisco (CA)

On-site
USD 130,000 - 195,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Deepgram • San Francisco (CA)

On-site
USD 150,000 - 230,000
Embedded AI Engineer: Edge Hardware & Kernels
Embedded AI Engineer: Edge Hardware & Kernels

Aimlroles • Northern (KY)

Hybrid
USD 120,000 - 190,000
Embedded AI Engineer: On-Device Speech & Edge Inference
Embedded AI Engineer: On-Device Speech & Edge Inference

Deepgram, Inc. • United States

On-site
USD 140,000 - 190,000