Get more replies from employers
Send a job-specific resume in minutes.
AWS is seeking a software engineer in the Neuron Compiler team to build the next-generation compiler translating PyTorch, TensorFlow, and JAX models for AWS Inferentia and Trainium servers in the cloud.
You will optimize compiler performance, collaborate with chip architects and ML apps teams, and contribute to open-source projects to advance ML workloads on AWS hardware.
Do you want to be part of the AI revolution? At AWS our vision is to make deep learning pervasive for everyday developers and to democratize access to AI hardware and software infrastructure. To deliver on that vision, we’ve created innovative software and hardware solutions that make it possible. AWS Neuron is the SDK that optimizes the performance of complex ML models executed on AWS Inferentia and Trainium, our custom chips designed to accelerate deep‑learning workloads.
This role is for a software engineer in the Compiler team for AWS Neuron. You will build next‑generation Neuron compiler that transforms ML models written in frameworks such as PyTorch, TensorFlow, and JAX to run on AWS Inferentia and Trainium based servers in the Amazon cloud. Your work will involve solving hard compiler optimization problems to achieve optimum performance for a variety of ML model families, including massive‑scale large language models like Llama and Deepseek, as well as stable diffusion, vision transformers, and multi‑model setups. You will deeply understand how these models work internally to inform compiler design decisions, communicate with internal and external stakeholders, and participate in pre‑silicon design to bring new products and features to market. The goal is to make the Neuron compiler highly performant and easy‑to‑use.
Key responsibilities include designing, implementing, testing, deploying, and maintaining innovative software solutions that improve the Neuron compiler’s performance, stability, and user interface. You will work closely with chip architects, runtime/OS engineers, scientists, and ML apps teams to deploy state‑of‑the‑art ML models on AWS accelerators with optimal cost/performance benefits. You will also engage with open‑source projects such as StableHLO, OpenXLA, and MLIR to pioneer advanced ML workload optimization on AWS hardware, build features that deliver great developer experiences, and create tools to analyze numerical errors and resolve compiler defects.
Experience in object‑oriented languages like C++/Java is a must. Experience with compilers or building ML models on accelerators (e.g., GPUs) is preferred but not required. Familiarity with OpenXLA, StableHLO, and MLIR is a bonus.
Amazon is an equal‑opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Benefits: The base salary range for this position is USD 165,200.00 – 223,600.00 annually. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D, optional Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more at https://amazon.jobs/en/benefits.