Get more replies from employers
Send a job-specific resume in minutes.
Amazon Web Services (AWS) is seeking an experienced software engineer to optimize collective operations and communication patterns for Trainium, enabling scalable AI compute across data centers.
In collaboration with hardware teams, you will co-optimize software and silicon for modern LLM training topologies, focusing on performance, reliability, and efficiency.
Optimize collective operations and communication patterns for AWS Trainium to scale AI compute across data centers. Collaborate with hardware teams to co‑optimize software and silicon for modern LLM training topologies.
Requires a Bachelor's degree in Computer Science and at least 5 years of experience in building complex software systems and architecture. Familiarity with collective communication algorithms or distributed training frameworks is preferred.
C++, Collective Communication Algorithms, Distributed Systems, DMA, Firmware, AI Training, LLM Topologies, Software Architecture, Performance Optimization, Neuron Explorer, Bus Bandwidth Optimization, System Design