Get more replies from employers
Send a job-specific resume in minutes.
Amazon is seeking a Senior Inference Engineer to own the end-to-end real-time multimodal inference stack. You will shape model architectures for servability, build real-time runtimes, and scale offline systems for RL and evaluation.
The role spans research-to-production, requiring deep knowledge of transformers, KV-cache, and GPU optimization. You will collaborate with scientists and hardware partners to deliver sub-second latency under concurrent load, while balancing performance and cost.
Amazon is seeking a Senior Inference Engineer to own the end-to-end real-time multimodal inference stack. You will shape model architectures for servability, build real-time runtimes, and scale offline systems for RL and evaluation.
The role spans research-to-production, requiring deep knowledge of transformers, KV-cache, and GPU optimization. You will collaborate with scientists and hardware partners to deliver sub-second latency under concurrent load, while balancing performance and cost.