Senior Real-Time Multimodal Inference Engineer

Amazon

Seattle (WA)

On-site

USD 167,000 - 226,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave
RSUs

Job summary

Amazon in Seattle is seeking a Senior Inference Engineer to own end-to-end real-time inference for multimodal models, spanning research to production. You will shape architectures for servability, build streaming runtimes, and develop offline systems for RL and post-training rollout.

You will collaborate with scientists and hardware partners to achieve sub-second latency on real-time workloads while controlling cost and ensuring scalability across distributed systems and multiple GPUs.

Qualifications

  • 3+ years building ML models for business use cases.
  • PhD or MS with 6+ years of applied research.
  • Experience programming in Java, C++, Python or related language.
  • Hands-on experience optimizing inference for neural models and profiling.

Responsibilities

  • Partner with researchers to surface latency, memory, and cost implications early.
  • Implement and optimize inference paths for large multimodal models with low latency.
  • Apply efficiency techniques like quantization and decoding trade-offs.
  • Develop high-performance kernels and integrate into production serving.
  • Profile end-to-end performance with Nsight Compute/Systems and roofline analysis.
  • Own real-time streaming serving path and meet sub-second latency under load.
  • Brandish continuous batching, scheduling, and preemption for throughput vs latency.
  • Customize production serving frameworks for real-time streaming generative models.
  • Implement multi-GPU inference and drive cost towards production baselines.
  • Publish latency, throughput, and cost benchmarks to gate deployment.

Skills

ML model building
Java/C++/Python
Inference optimization
GPU performance
Production scale delivery

Education

PhD or MS + 6y

Tools

vLLM
PyTorch
TensorRT-LLM
NCCL

Job description

Amazon in Seattle is seeking a Senior Inference Engineer to own end-to-end real-time inference for multimodal models, spanning research to production. You will shape architectures for servability, build streaming runtimes, and develop offline systems for RL and post-training rollout.

You will collaborate with scientists and hardware partners to achieve sub-second latency on real-time workloads while controlling cost and ensuring scalability across distributed systems and multiple GPUs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA), Northern (KY)

Hybrid
USD 167,000 - 226,000
Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Senior Real-Time Multimodal AI Scientist
Senior Real-Time Multimodal AI Scientist

Amazon Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 192,000 - 260,000
Senior ML Inference Engineer — Open-Source LLM Serving
Senior ML Inference Engineer — Open-Source LLM Serving

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
Senior Real-Time Multimodal Conversational AI Scientist
Senior Real-Time Multimodal Conversational AI Scientist

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Senior Applied Scientist, Real-Time Multimodal AI
Senior Applied Scientist, Real-Time Multimodal AI

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Lead Scientist Real-Time Multimodal Conversational AI
Lead Scientist Real-Time Multimodal Conversational AI

Amazon • Boston (MA)

On-site
USD 167,000 - 226,000
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+2
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Boston (MA), Northern (KY)

On-site
USD 167,000 - 226,000
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options