Production ML Engineer: Turn Research into Scalable Models

Deepgram, Inc.

Northern (KY)

Hybrid

USD 140,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Deepgram, Inc. is building the real-time speech AI stack for a trillion-dollar market. We are seeking an Applied ML Engineer to own the research-to-production pipeline, collaborating with researchers to deploy models at scale across our hybrid infrastructure.

You will design release gates, optimize inference, and build tooling to accelerate production-ready AI. Strong Python, PyTorch, and systems experience are essential, with a track record shipping ML at scale.

Qualifications

  • Strong software engineering fundamentals, with proficiency in Python and production-quality ML code.
  • Hands-on experience taking ML models from research or prototype stage into production at scale.
  • Familiarity with serving and inference optimization for production workloads.

Responsibilities

  • Own the research-to-production pipeline: turn research checkpoints into production models and define the path to deployment.
  • Partner with researchers to productionize new models, translating experimental code into robust, reproducible workflows.
  • Build tooling and abstractions to move models through training, evaluation, packaging, and deployment with minimal friction.
  • Design and own model release gates with automated evaluation, regression detection, and quality checks for shipping models.
  • Optimize models and serving for production: efficient inference, batching, memory and latency tuning.
  • Strengthen the build and delivery layer for models on our hybrid infrastructure (GPU data centers and cloud).
  • Establish benchmarking and validation across development to production to catch regressions early.
  • Build feedback loops to surface what's working and drive the next iteration.

Skills

Python
PyTorch
ML pipelines
Distributed systems
Production ML

Tools

CI/CD for ML
Model packaging

Job description

Deepgram, Inc. is building the real-time speech AI stack for a trillion-dollar market. We are seeking an Applied ML Engineer to own the research-to-production pipeline, collaborating with researchers to deploy models at scale across our hybrid infrastructure.

You will design release gates, optimize inference, and build tooling to accelerate production-ready AI. Strong Python, PyTorch, and systems experience are essential, with a track record shipping ML at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied ML Engineer: From Research to Production
Applied ML Engineer: From Research to Production

deepgram • United States

On-site
USD 140,000 - 180,000
ML Engineer: Research-to-Production Pipelines
ML Engineer: Research-to-Production Pipelines

Deepgram • United States

On-site
USD 180,000 - 240,000
Staff ML Engineer – Scalable Speech Systems
Staff ML Engineer – Scalable Speech Systems

Deepgram • Ann Arbor (MI)

On-site
USD 150,000 - 230,000
Applied ML Engineer
Applied ML Engineer

deepgram • United States

On-site
USD 140,000 - 180,000
Applied ML Engineer
Applied ML Engineer

Deepgram, Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Research Engineer, Machine Learning Systems
Research Engineer, Machine Learning Systems

Deepgram • Ann Arbor (MI)

On-site
USD 150,000 - 230,000
Senior ML Engineer — Production-Ready Voice AI (Hybrid)
Senior ML Engineer — Production-Ready Voice AI (Hybrid)

Modulate • Somerville (MA)

Hybrid
USD 170,000 - 200,000
Competitive salary + equity
Full health, dental, and vision
Flexible PTO
+5
Engineering Manager, ML Research & Production Equity
Engineering Manager, ML Research & Production Equity

Speedrun Talent Network • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 430,000
Medical, Dental, Vision
Life Insurance
Disability Benefits
+6
Senior Technical Program Manager (Engineering) - AI Tooling & Systems
Senior Technical Program Manager (Engineering) - AI Tooling & Systems

Deepgram • United States

On-site
USD 180,000 - 240,000
Research Engineer: Real-Time Voice ML & Production Systems
Research Engineer: Real-Time Voice ML & Production Systems

Sesame • Bellevue (WA)

On-site
USD 190,000 - 320,000
401(k) max match
Health, vision & dental benefits
Unlimited PTO"
+3