Senior Real-Time Multimodal Inference Engineer

Amazon

Boston (MA)

On-site

USD 168,000 - 227,000

Full time

27 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon is seeking a Senior Inference Engineer to own end-to-end inference for real-time multimodal conversational AI in a full-stack role. You will shape model architectures for servability, build real-time runtimes, and create offline systems for training and reinforcement learning.

You will work across research, engineering, and hardware teams to ensure sub-second latency while managing cost. You will co-design architectures with scientists, optimize KV-cache and attention, and drive efficient

Qualifications

  • 5+ years of professional software development experience.
  • 5+ years of programming with at least one programming language.
  • 4+ years of leading design or architecture of new and existing systems.
  • Bachelor's degree in computer science or equivalent.
  • Experience as a mentor or tech lead.
  • 2+ years of hands-on experience optimizing inference for neural models.
  • Strong understanding of deep learning architectures and multimodal domains.
  • Production experience delivering latency-constrained, real-time inference systems.

Responsibilities

  • Partner with research scientists to surface latency, memory, and cost implications of architecture choices.
  • Implement and optimize the inference path for large-scale multimodal models.
  • Apply efficiency techniques across the stack and quantify latency trade-offs.
  • Develop high-performance kernels for critical operations and integrate into production serving.
  • Profile end-to-end performance and identify bottlenecks in large-scale inference workloads.
  • Own the real-time serving path for streaming multimodal conversational AI.
  • Build and tune batching, scheduling, and preemption for throughput vs latency.
  • Customize production serving frameworks for real-time streaming generative models.
  • Implement multi-GPU inference and drive costs toward production baselines.
  • Establish latency, throughput, and cost benchmarks for deployment.

Skills

Inference engineering
Latency optimization
GPU optimization
Python
C++

Education

Bachelor's degree in CS or equivalent

Tools

Nsight Compute
PyTorch
vLLM
CUDA

Job description

Amazon is seeking a Senior Inference Engineer to own end-to-end inference for real-time multimodal conversational AI in a full-stack role. You will shape model architectures for servability, build real-time runtimes, and create offline systems for training and reinforcement learning.

You will work across research, engineering, and hardware teams to ensure sub-second latency while managing cost. You will co-design architectures with scientists, optimize KV-cache and attention, and drive efficient

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Socket.dev • Boston (MA)

On-site
USD 168,000 - 227,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Sunnyvale (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Senior Real-Time Multimodal Conversational AI Scientist
Senior Real-Time Multimodal Conversational AI Scientist

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+1
Senior Real-Time Multimodal AI Scientist
Senior Real-Time Multimodal AI Scientist

Amazon Science • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Realtime Multimodal Inference Architect
Realtime Multimodal Inference Architect

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Boston (MA)

On-site
USD 168,000 - 227,000
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Socket.dev • Boston (MA)

On-site
USD 168,000 - 227,000
Senior Real-Time Conversational AI Scientist (AGI)
Senior Real-Time Conversational AI Scientist (AGI)

Amazon • Boston (MA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+1
Principal Scientist, Real-Time Multimodal Conversational AI
Principal Scientist, Real-Time Multimodal Conversational AI

Amazon • Bellevue (WA)

On-site
USD 199,000 - 269,000
Principal Scientist, Real-Time Multimodal Conversational AI
Principal Scientist, Real-Time Multimodal Conversational AI

Amazon • Boston (MA)

On-site
USD 199,000 - 269,000
Health insurance
RSUs
401(k) matching
+1