Real-Time Multimodal Inference Engineer

Amazon

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

16 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon in Seattle seeks an Inference Engineer to own components of real-time multimodal inference, from research to production, ensuring sub-second latency within cost constraints.

You will collaborate with scientists and senior engineers to optimize the inference stack, develop high-performance kernels, and shape model architectures to be servable. This role blends research and production in a fast-paced environment.

Qualifications

  • Master's degree or equivalent.
  • 3+ years non-internship software development experience.
  • 3+ years programming in at least one language.
  • 2+ years design/architecture experience.
  • 1+ years optimizing neural inference.
  • Solid understanding of transformers and speech/audio applications.
  • Experience with latency-constrained real-time inference.
  • Experience with GPU performance optimization.

Responsibilities

  • Partner with researchers to help make model architectures servable and surface latency and cost implications.
  • Implement and optimize parts of the inference path for large-scale multimodal models and KV-cache.
  • Apply efficiency techniques across the stack and measure latency trade-offs.
  • Develop high-performance kernels and integrate into production serving.
  • Profile end-to-end performance using Nsight Compute/Systems and roofline analysis.
  • Contribute to real-time streaming serving paths and latency budgets.
  • Build offline inference systems for RL/RLHF/RLAIF and train/serve parity.

Skills

Software dev
Programming languages
Systems design
GPU optimization
Real-time inference
Ownership

Education

Master's degree or equivalent
Bachelor's degree in CS or equivalent

Tools

vLLM
TensorRT-LLM
CUTLASS
Triton
FlashAttention
KV-cache

Job description

Amazon in Seattle seeks an Inference Engineer to own components of real-time multimodal inference, from research to production, ensuring sub-second latency within cost constraints.

You will collaborate with scientists and senior engineers to optimize the inference stack, develop high-performance kernels, and shape model architectures to be servable. This role blends research and production in a fast-paced environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+2
Real-Time Multimodal Inference Engineer
Real-Time Multimodal Inference Engineer

Amazon • Sunnyvale (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon Inc. • Boston (MA)

On-site
USD 167,000 - 226,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA), Northern (KY)

Hybrid
USD 167,000 - 226,000
Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Multimodal Inference Engineer
Multimodal Inference Engineer

OpenAI • United States

Remote
USD 325,000 - 490,000
Medical insurance
Mental health support
401(k) plan with 50% matching
+3
Senior ML Inference Engineer – High-Performance Serving
Senior ML Inference Engineer – High-Performance Serving

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
Inference Engineer, AGI
Inference Engineer, AGI

Amazon • Sunnyvale (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
ML Inference Systems Engineer — High-Performance Serving
ML Inference Systems Engineer — High-Performance Serving

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
RSUs
401(k) matching
Paid time off
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave