Edge ML Engineer: Port & Deploy Models to Edge Devices

Precision Labs

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Deepgram is seeking an Applied ML Engineer for the Partner Platform Engineering team to port and deploy speech models to edge hardware, including non-NVIDIA platforms. You will adapt models, optimize quantization and precision, and ensure latency and accuracy on real devices.

This role emphasizes production deployment over research, requiring collaboration with Embedded AI Engineers and platform vendors. Seniority will be set by experience and needs.

Qualifications

  • Hands-on experience deploying ML models to edge or non-NVIDIA hardware in production.
  • Experience with quantization and precision tradeoffs (INT8, FP16, mixed precision, calibration).
  • Proficiency in Python and PyTorch with production-quality practices.
  • Ability to modify model graphs, swap operators, and adjust architecture without breaking accuracy.
  • Experience with at least one edge or vendor inference runtime and its conversion toolchain.

Responsibilities

  • Port Deepgram speech models to non-NVIDIA and edge platforms with minimal modification to run within the target runtime.
  • Own serving-side model decisions for edge targets: quantization, precision, operator substitutions, and graph rewrites.
  • Validate ports on real hardware: accuracy, latency, throughput, and memory benchmarks; catch regressions early.
  • Build the deployment path for edge targets: packaging, conversion pipelines, versioning, and automated delivery.
  • Collaborate with Embedded AI Engineers for custom kernels and port adaptation.
  • Coordinate with platform and silicon vendors on runtimes and toolchains to consistent model formats.
  • Provide feedback to Research/Impeller to simplify future ports.
  • Assist with production concerns like automated deployment and fleet observability.

Skills

Edge deployment
Quantization & precision
Python
PyTorch
Model graphs & operator rewriting

Tools

ONNX Runtime
TFLite
ExecuTorch
OpenVINO
Qualcomm AI Engine
Vendor NPU SDK

Job description

Deepgram is seeking an Applied ML Engineer for the Partner Platform Engineering team to port and deploy speech models to edge hardware, including non-NVIDIA platforms. You will adapt models, optimize quantization and precision, and ensure latency and accuracy on real devices.

This role emphasizes production deployment over research, requiring collaboration with Embedded AI Engineers and platform vendors. Seniority will be set by experience and needs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Edge ML Engineer: Port Speech Models to Diverse Hardware
Edge ML Engineer: Port Speech Models to Diverse Hardware

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Precision Labs • Northern (KY)

Hybrid
USD 150,000 - 210,000
Applied ML Engineer - Edge Devices
Applied ML Engineer - Edge Devices

Madrona Venture Labs • United States

On-site
USD 140,000 - 210,000
Edge AI Engineer — Real‑Time, On‑Device Speech Models
Edge AI Engineer — Real‑Time, On‑Device Speech Models

Deepgram • United States

On-site
USD 140,000 - 190,000
Applied ML Engineer: From Research to Production
Applied ML Engineer: From Research to Production

deepgram • United States

On-site
USD 140,000 - 180,000
Embedded AI Engineer: On-Device Speech & Edge Inference
Embedded AI Engineer: On-Device Speech & Edge Inference

Deepgram, Inc. • United States

On-site
USD 140,000 - 190,000
Edge ML Engineer: Real-Time Speech & Language for Autonomy
Edge ML Engineer: Real-Time Speech & Language for Autonomy

Instant Teams • New Haven (CT)

On-site
USD 150,000 - 215,000
Equity stake
100% Remote (US)
Comprehensive Health Coverage
+3
Production ML Engineer: Turn Research into Scalable Models
Production ML Engineer: Turn Research into Scalable Models

Deepgram, Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Edge AI Engineer: On-Device ML for Production (Remote)
Edge AI Engineer: On-Device ML for Production (Remote)

Bright Vision Technologies • Columbus (OH), Dublin (OH)

Remote
USD 100,000 - 150,000
Embedded AI Engineer, On-Device Models
Embedded AI Engineer, On-Device Models

Deepgram • United States

On-site
USD 140,000 - 190,000