Get more replies from employers
Send a job-specific resume in minutes.
Amazon is seeking a Senior Inference Engineer to own real-time multimodal inference across research to production, shaping model architectures for servable deployment, building the low-latency runtime, and supporting offline training systems.
You will collaborate with scientists and hardware partners to ensure models run under strict latency budgets, own end-to-end inference stack, and explore cross-cutting optimizations across the architecture, streaming serving, and RL/evaluation pipelines.
Amazon is seeking a Senior Inference Engineer to own real-time multimodal inference across research to production, shaping model architectures for servable deployment, building the low-latency runtime, and supporting offline training systems.
You will collaborate with scientists and hardware partners to ensure models run under strict latency budgets, own end-to-end inference stack, and explore cross-cutting optimizations across the architecture, streaming serving, and RL/evaluation pipelines.