Stand out for this role — generate a tailored resume and cover letter in about a minute.
Arm in Seattle is seeking a Software Engineer for the AI Inference Runtime team to set technical direction for distributed inference runtime powering SOTA AI models. You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed execution, kernel development and optimization, shaping how efficiently models use compute.
You will partner with AI Infrastructure, compute and product teams to raise performance and energy efficiency of Arm’s AI platform,
Arm in Seattle is seeking a Software Engineer for the AI Inference Runtime team to set technical direction for distributed inference runtime powering SOTA AI models. You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed execution, kernel development and optimization, shaping how efficiently models use compute.
You will partner with AI Infrastructure, compute and product teams to raise performance and energy efficiency of Arm’s AI platform,