Stand out for this role — generate a tailored resume and cover letter in about a minute.
Amazon in Seattle is seeking a Senior Inference Engineer to own end-to-end real-time inference for multimodal models, spanning research to production. You will shape architectures for servability, build streaming runtimes, and develop offline systems for RL and post-training rollout.
You will collaborate with scientists and hardware partners to achieve sub-second latency on real-time workloads while controlling cost and ensuring scalability across distributed systems and multiple GPUs.
Amazon in Seattle is seeking a Senior Inference Engineer to own end-to-end real-time inference for multimodal models, spanning research to production. You will shape architectures for servability, build streaming runtimes, and develop offline systems for RL and post-training rollout.
You will collaborate with scientists and hardware partners to achieve sub-second latency on real-time workloads while controlling cost and ensuring scalability across distributed systems and multiple GPUs.