A complete application in a minute — tailored resume and cover letter, ready to send.
Amazon AWS Neuron is seeking a senior software engineer for the Machine Learning Inference Applications team in Seattle. You will develop and optimize core building blocks of LLM inference on Neuron chips, including attention, MLP, quantization, speculative decoding, and Mixture of Experts.
Responsibilities include applying the latest research in LLM optimization to extract top performance from open source and internal models, and working across teams to deliver scalable, high-performance
Amazon AWS Neuron is seeking a senior software engineer for the Machine Learning Inference Applications team in Seattle. You will develop and optimize core building blocks of LLM inference on Neuron chips, including attention, MLP, quantization, speculative decoding, and Mixture of Experts.
Responsibilities include applying the latest research in LLM optimization to extract top performance from open source and internal models, and working across teams to deliver scalable, high-performance