Get more replies from employers
Send a job-specific resume in minutes.
Amazon is seeking a Senior Inference Engineer to own end-to-end inference for real-time multimodal conversational AI in a full-stack role. You will shape model architectures for servability, build real-time runtimes, and create offline systems for training and reinforcement learning.
You will work across research, engineering, and hardware teams to ensure sub-second latency while managing cost. You will co-design architectures with scientists, optimize KV-cache and attention, and drive efficient
Amazon is seeking a Senior Inference Engineer to own end-to-end inference for real-time multimodal conversational AI in a full-stack role. You will shape model architectures for servability, build real-time runtimes, and create offline systems for training and reinforcement learning.
You will work across research, engineering, and hardware teams to ensure sub-second latency while managing cost. You will co-design architectures with scientists, optimize KV-cache and attention, and drive efficient