Senior Real-Time Multimodal Inference Engineer

Amazon

Boston, Northern (MA, KY)

Hybrid

USD 167,000 - 226,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon.com Services LLC is seeking a Senior Inference Engineer to own the inference stack for real-time multimodal conversational AI. You will shape model architectures for servable deployment, build real-time runtimes within tight latency budgets, and develop offline systems for training and RL integration.

You will collaborate with scientists and hardware teams to ensure fast, cost-effective inference at scale, across training, evaluation, and deployment pathways.

Qualifications

  • 3+ years of building ML models for business applications.
  • PhD, or Master's with 6+ years of applied research experience.
  • Experience with Java, C++, Python or related language.
  • 2+ years hands-on experience optimizing inference for neural models.
  • Strong understanding of transformers, attention, and autoregressive decoding.
  • Production track record delivering latency-constrained real-time inference.

Responsibilities

  • Own the real-time streaming inference path for multimodal models.
  • Develop and optimize the end-to-end inference stack from research to production.
  • Profile and optimize latency, memory, and cost across the stack.
  • Collaborate with scientists and hardware partners to meet latency budgets.
  • Establish performance benchmarks and publish operational metrics.

Skills

Latency optimization
Real-time inference
Multimodal models
System design
Collaboration

Education

PhD or Master's + research

Tools

Java
C++
Python
Nsight Compute
TensorRT-LLM
vLLM
CUDA

Job description

Amazon.com Services LLC is seeking a Senior Inference Engineer to own the inference stack for real-time multimodal conversational AI. You will shape model architectures for servable deployment, build real-time runtimes within tight latency budgets, and develop offline systems for training and RL integration.

You will collaborate with scientists and hardware teams to ensure fast, cost-effective inference at scale, across training, evaluation, and deployment pathways.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Senior Real-Time Multimodal AI Scientist
Senior Real-Time Multimodal AI Scientist

Amazon Science • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Senior Applied Scientist, Real-Time Multimodal AI
Senior Applied Scientist, Real-Time Multimodal AI

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Lead Real-Time Multimodal AI Scientist, Conversational/AGI
Lead Real-Time Multimodal AI Scientist, Conversational/AGI

Amazon • Bellevue (WA)

On-site
USD 180,000 - 260,000
Realtime Multimodal Inference Architect
Realtime Multimodal Inference Architect

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Lead Real-Time Multimodal AI Scientist
Lead Real-Time Multimodal AI Scientist

Amazon • Sunnyvale (CA)

On-site
USD 229,000 - 309,000
Lead Scientist - Real-Time Multimodal Conversational AI
Lead Scientist - Real-Time Multimodal Conversational AI

Amazon.com Services LLC • Seattle (WA)

On-site
USD 210,000 - 310,000
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Boston (MA), Northern (KY)

On-site
USD 167,000 - 226,000
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Senior Real-Time Conversational AI Scientist (AGI)
Senior Real-Time Conversational AI Scientist (AGI)

Amazon • Boston (MA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+1