Real-Time Multimodal Inference Engineer

Amazon

Sunnyvale (CA)

On-site

USD 165,000 - 224,000

Full time

23 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon is seeking an Inference Engineer in Sunnyvale to own components of the inference stack and drive their execution with guidance from scientists and senior engineers. You will shape model architectures for servability, build real-time runtimes, and develop offline systems for training and reinforcement learning."

You will collaborate across teams to ensure neural models run within hard latency budgets on production hardware, balancing throughput and cost at scale.

Qualifications

  • Master's degree or higher in CS or related field.
  • 3+ years of professional software development experience.
  • 2+ years in design/architecture of new/existing systems.
  • 1+ years optimizing inference for neural models.
  • Deep learning architectures (transformers) and multimodal models.

Responsibilities

  • Help make model architectures servable with latency, memory, and cost focus.
  • Implement and optimize inference paths for large-scale multimodal models.
  • Develop high-performance kernels for critical operations in production serving.
  • Profile end-to-end performance and identify bottlenecks in large-scale inference workloads.
  • Support real-time streaming serving with sub-second latency under concurrent load.

Skills

Latency optimization
GPU performance
Distributed systems
Real-time inference
Python

Education

Master's degree
Bachelor's degree in CS

Tools

vLLM
TensorRT-LLM
CUDA
Triton

Job description

Amazon is seeking an Inference Engineer in Sunnyvale to own components of the inference stack and drive their execution with guidance from scientists and senior engineers. You will shape model architectures for servability, build real-time runtimes, and develop offline systems for training and reinforcement learning."

You will collaborate across teams to ensure neural models run within hard latency budgets on production hardware, balancing throughput and cost at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Real-Time Multimodal Inference Engineer
Real-Time Multimodal Inference Engineer

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA), Northern (KY)

Hybrid
USD 167,000 - 226,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+2
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon Inc. • Boston (MA)

On-site
USD 167,000 - 226,000
Multimodal Inference Engineer
Multimodal Inference Engineer

OpenAI • United States

Remote
USD 325,000 - 490,000
Medical insurance
Mental health support
401(k) plan with 50% matching
+3
Senior ML Inference Engineer – High-Performance Serving
Senior ML Inference Engineer – High-Performance Serving

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
AI/ML Systems Engineer - High-Performance Inference
AI/ML Systems Engineer - High-Performance Inference

Amazon Inc. • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
ML Inference Systems Engineer — High-Performance Serving
ML Inference Systems Engineer — High-Performance Serving

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
RSUs
401(k) matching
Paid time off