Edge ML Engineer: Port & Optimize Models for Edge

Aimlroles

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Deepgram is seeking an Applied ML Engineer for the Partner Platform Engineering team to port speech models to edge and non-NVIDIA hardware. You will adapt models, validate on real devices, and own the deployment pipeline.

You’ll work with embedded engineers, quantify latency, and optimize for constrained devices, shipping models to customers rather than publishing research, with a senior-to-staff level title depending on experience.

Qualifications

  • Experience deploying ML models to edge or non-GPU hardware in production.
  • Strong Python and PyTorch, with production-quality engineering habits.
  • Experience with quantization and precision tradeoffs (INT8, FP16, mixed precision).
  • Ability to modify models to fit platforms by swapping operators and adjusting architecture.

Responsibilities

  • Port Deepgram speech models to non-NVIDIA and edge platforms.
  • Own serving-side model decisions for edge targets: quantization and precision choices.
  • Validate port performance on real hardware and catch regressions.
  • Build deployment paths for edge targets: packaging and versioning.
  • Collaborate with Embedded AI Engineers on custom kernels.
  • Partner with platform vendors on runtimes and toolchains.
  • Feed edge constraints back to Research and Impeller.
  • Take on adjacent production concerns at the edge.

Skills

Python
PyTorch
Edge deployment
Quantization
Model porting
CUDA
Metal
NEON
ONNX Runtime
TFLite
ExecuTorch
OpenVINO
Qualcomm AI Engine
Vendor NPU SDK

Tools

ONNX Runtime
TFLite
ExecuTorch
OpenVINO
Qualcomm AI Engine
Vendor NPU SDK

Job description

Deepgram is seeking an Applied ML Engineer for the Partner Platform Engineering team to port speech models to edge and non-NVIDIA hardware. You will adapt models, validate on real devices, and own the deployment pipeline.

You’ll work with embedded engineers, quantify latency, and optimize for constrained devices, shipping models to customers rather than publishing research, with a senior-to-staff level title depending on experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Edge ML Engineer: Port Models to Non-NVIDIA & Devices
Edge ML Engineer: Port Models to Non-NVIDIA & Devices

Deepgram • San Francisco (CA)

On-site
USD 150,000 - 230,000
Edge ML Engineer: Port & Deploy Models to Edge Devices
Edge ML Engineer: Port & Deploy Models to Edge Devices

Precision Labs • Northern (KY)

Hybrid
USD 150,000 - 210,000
Edge ML Engineer: Port Speech Models to Diverse Hardware
Edge ML Engineer: Port Speech Models to Diverse Hardware

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Edge ML Engineer: Port Models to Non-NVIDIA & ARM
Edge ML Engineer: Port Models to Non-NVIDIA & ARM

Deepgram • United States

Remote
USD 150,000 - 230,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Precision Labs • Northern (KY)

Hybrid
USD 150,000 - 210,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Deepgram • San Francisco (CA)

On-site
USD 150,000 - 230,000
On-Device AI Engineer: Edge Model Optimization
On-Device AI Engineer: Edge Model Optimization

Deepgram • United States

Remote
USD 140,000 - 210,000
Embedded AI Engineer: On-Device Models & Edge Kernels
Embedded AI Engineer: On-Device Models & Edge Kernels

Deepgram • San Francisco (CA)

On-site
USD 130,000 - 195,000
Embedded AI Engineer: On-Device Speech & Edge Inference
Embedded AI Engineer: On-Device Speech & Edge Inference

Deepgram, Inc. • United States

On-site
USD 140,000 - 190,000