Senior Real-Time Multimodal Inference Engineer

Socket.dev

Boston (MA)

On-site

USD 168,000 - 227,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon is seeking a Senior Inference Engineer to own real-time multimodal inference across research to production, shaping model architectures for servable deployment, building the low-latency runtime, and supporting offline training systems.

You will collaborate with scientists and hardware partners to ensure models run under strict latency budgets, own end-to-end inference stack, and explore cross-cutting optimizations across the architecture, streaming serving, and RL/evaluation pipelines.

Qualifications

  • 5+ years of non-internship professional software development experience.
  • 4+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems.
  • Bachelor's degree in computer science or equivalent.
  • Experience as a mentor, tech lead or leading an engineering team.
  • 2+ years of hands‑on experience optimizing inference for neural models — not just using inference frameworks, but profiling and improving them.
  • Strong understanding of deep learning architectures (transformers, attention mechanisms, autoregressive decoding) and their application to speech/audio or other multimodal domains.
  • Production track record delivering latency-constrained, real-time inference systems under concurrent load.
  • Experience with GPU performance optimization — memory hierarchy, occupancy, KV-cache management, and the accelerator programming model.
  • Demonstrated ownership of a technical area — driving execution for a workstream and collaborating effectively across scientists and engineers.

Responsibilities

  • Partner with research scientists to make model architectures servable from inception — surfacing the latency, memory, and cost implications of architecture choices before they are locked in
  • Implement and optimize the inference path for large-scale multimodal models — attention and KV-cache mechanisms, multimodal/autoregressive decoding, and the compute primitives on the critical path
  • Apply efficiency techniques across the stack — quantization (per-tensor/per-channel/per-group, INT8/FP8/BF16), speculative decoding, operator fusion, and paged KV-cache — and quantify their quality/latency trade-offs
  • Develop and tune high-performance kernels for critical operations where off-the-shelf implementations leave performance on the table, integrating them into production serving with minimal overhead
  • Profile end-to-end performance with tools such as Nsight Compute/Systems and roofline analysis to identify and eliminate bottlenecks in large-scale inference workloads
  • Own the real-time serving path for streaming multimodal conversational AI, meeting sub-second, streaming latency budgets under concurrent session load
  • Build and tune continuous batching, scheduling, and preemption to balance throughput against per-request latency SLAs for interactive workloads
  • Customize production serving frameworks (e.g., vLLM, PyTorch) for real-time streaming generative models that fall outside standard LLM serving patterns — sustained low-latency output under concurrent session load
  • Implement multi-GPU inference (tensor parallelism, collective communication) for latency-critical paths, and drive cost toward parity with existing production baselines
  • Establish latency, throughput, and cost benchmarking, and publish the operational metrics that gate deployment
  • Build offline inference systems behind post-training — high-throughput rollout generation and reward-model serving for reinforcement learning (RL/RLHF/RLAIF)
  • Ensure train/serve consistency — that the inference path used in RL and evaluation faithfully matches production online behavior
  • Work with the evaluation team to enable offline inference that captures the quality dimensions unique to real-time conversation — latency sensitivity, audio quality, and interaction naturalness

Skills

5+ years of non-internshipprofessional
Strong understanding of deep learning
4+ years leading design/architecture
2+ years optimizing inference for ней-

Education

Bachelor's degree in computer science or equivalent

Tools

vLLM
TensorRT-LLM
CUTLASS
FlashAttention
Triton

Job description

Amazon is seeking a Senior Inference Engineer to own real-time multimodal inference across research to production, shaping model architectures for servable deployment, building the low-latency runtime, and supporting offline training systems.

You will collaborate with scientists and hardware partners to ensure models run under strict latency budgets, own end-to-end inference stack, and explore cross-cutting optimizations across the architecture, streaming serving, and RL/evaluation pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Sunnyvale (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Real-Time Multimodal Inference Architect
Real-Time Multimodal Inference Architect

Amazon • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
RSUs
401(k) matching
+2
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA)

On-site
USD 168,000 - 227,000
Realtime Multimodal Inference Architect
Realtime Multimodal Inference Architect

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Boston (MA)

On-site
USD 168,000 - 227,000
Senior Real-Time Multimodal Conversational AI Scientist
Senior Real-Time Multimodal Conversational AI Scientist

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+1
Senior Real-Time Multimodal AI Scientist
Senior Real-Time Multimodal AI Scientist

Amazon Science • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Socket.dev • Boston (MA)

On-site
USD 168,000 - 227,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Sunnyvale (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1