ML Engineer: Research-to-Production Pipelines

Deepgram

United States

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Deepgram is seeking an Applied ML Engineer to own the research‑to‑production pipeline and streamline model delivery at scale. You’ll work with researchers to productionize novel models, build robust tooling, and ensure repeatable, observable deployments.

You’ll be responsible for designing release gates, optimizing inference, and delivering scalable ML systems across our hybrid infrastructure.

Qualifications

  • Must have production‑quality ML code in Python.
  • Experience shipping ML models from research to production at scale.
  • Strong knowledge of PyTorch and the modern DL stack.
  • Experience building ML pipelines and developer tooling (training orchestration, evaluation harnesses, packaging, deployment, or CI/CD).
  • Familiarity with inference optimization (latency, batching, throughput) for production.
  • Comfort operating across distributed systems and GPU compute (cloud and on‑premise).

Responsibilities

  • Own the research‑to‑production pipeline: turn research checkpoints into production models and define repeatable deployment paths.
  • Collaborate with researchers to productionize new models with robust, reproducible workflows.
  • Build tooling and abstractions to move models through training, evaluation, packaging, and deployment with minimal friction.
  • Design and own model release gates with automated evaluation, regression checks, and timing decisions.
  • Optimize models and serving for production: reduce latency, improve throughput, and enable cost‑efficient inference.
  • Strengthen the build and delivery layer for models on hybrid GPU/cloud infrastructure.
  • Establish benchmarking and validation that catch regressions from development to production.
  • Create a feedback loop to surface model behavior and accelerate next iterations with research.

Skills

Python
Distributed systems
PyTorch
ML pipelines
Inference optimization
Production ML
Research-to-prod

Tools

PyTorch

Job description

Deepgram is seeking an Applied ML Engineer to own the research‑to‑production pipeline and streamline model delivery at scale. You’ll work with researchers to productionize novel models, build robust tooling, and ensure repeatable, observable deployments.

You’ll be responsible for designing release gates, optimizing inference, and delivering scalable ML systems across our hybrid infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied ML Engineer: From Research to Production
Applied ML Engineer: From Research to Production

deepgram • United States

On-site
USD 140,000 - 180,000
Production ML Engineer: Turn Research into Scalable Models
Production ML Engineer: Turn Research into Scalable Models

Deepgram, Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Applied ML Engineer
Applied ML Engineer

Deepgram, Inc. • Northern (KY)

On-site
USD 140,000 - 190,000
Senior AI Tools & Systems PM — ML Infra & Serving
Senior AI Tools & Systems PM — ML Infra & Serving

Deepgram, Inc. • United States

On-site
USD 150,000 - 230,000
Senior AI Infrastructure & Tooling Programs Lead
Senior AI Infrastructure & Tooling Programs Lead

Apply • United States

On-site
USD 150,000 - 190,000
Applied ML Engineer
Applied ML Engineer

deepgram • United States

On-site
USD 140,000 - 180,000
ML Engineer I: Production Models & Data Pipelines
ML Engineer I: Production Models & Data Pipelines

Compunnel, Inc. • San Francisco (CA)

On-site
USD 90,000 - 130,000
Senior ML Infrastructure TPM — AI Tooling & Systems
Senior ML Infrastructure TPM — AI Tooling & Systems

Deepgram • United States

Remote
USD 140,000 - 190,000
Senior ML Systems Engineer: Production Pipelines
Senior ML Systems Engineer: Production Pipelines

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI Infra & Tooling TPM – Real‑Time ML at Scale
Senior AI Infra & Tooling TPM – Real‑Time ML at Scale

Madrona Venture Labs • United States

On-site
USD 150,000 - 230,000