Job Overview
Do you want to build the backbone of Generative AI cloud at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price‑performance improvements for multi‑billion‑variable LLMs? Join us in designing, delivering, and operating AWS cloud offerings that enable high‑performance, scalable AI/ML and HPC workloads.
Key Responsibilities
You will work with engineers across the company to deliver the next‑generation AWS platforms. Your responsibilities include:
- Solving complex architectural problems that may not be defined before hand.
- Owning the team’s systems, proactively identifying deficiencies, writing tactical code to solve issues before they impact customers, and scaling solutions.
- Decomposing large server system testability, reliability, and diagnostics problems into manageable tasks or features, then leading delivery with other team members.
- Utilizing hardware, software, system design, x86 (and ARM), GPU/FPGA knowledge, and modern storage, networking, and memory technologies to diagnose and improve systems.
- Creating automation via agentic workflows, developing AI‑driven tools and workflows, and contributing to AI transformation.
- Collaborating with SDEs, SDETs, TPMs, managers, and principals across AWS to drive high quality and reliability into future designs and accelerator server solutions.
Basic Qualifications
- 4+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, or Ruby.
- 3+ years of professional software development experience (non‑internship).
- 3+ years of designing or architecting new and existing systems, focusing on reliability and scaling.
- Experience with x86 architecture, as well as ARM and GPU/FPGA devices.
- Knowledge of modern technology devices in storage, network, memory, and interface standards (I2C, IPMI, SPI, PCIe).
- BS degree in Computer Science, Computer Engineering, or related technical degree, or equivalent work experience.
Preferred Qualifications
- 7+ years or more of experience in software development, systems development, SRE, or resilience engineering.
- 7+ years of SysDE or equivalent experience.
- 7+ years of server systems debugging experience, including root cause analysis of complex server platforms.
- Experience in improving durability, security, availability, and scalability of systems through exploration, diagnosis, and remediation.
- Linux kernel and user‑space driver experience for PCIe and external devices.
- System thinking: diagnosing interactions between discrete components of a server system and driving product improvements.
- Strong focus on reliability, scale, and diagnostics; developing tactical and strategic tools using Python, Go, or C/C++.
- Solid understanding of OS internals, including network and storage subsystems.
- Master’s degree in Electrical Engineering, Computer Engineering, or related field.
- Experience validating hardware, software, firmware, and drivers, and implementing test plans.
- Experience with server validation, testing, root‑cause analysis, and coverage analysis.
- Excellent diagnostics tools development experience with Python, Go, or C/C++ in a fast‑paced environment.
- Extensive Linux knowledge.
Equal Employment Opportunity Statement
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Location and Compensation
USA, CA, Cupertino – 173,900.00 – 235,200.00 USD annually
USA, TX, Austin – 151,200.00 – 204,600.00 USD annually
USA, WA, Seattle – 151,200.00 – 204,600.00 USD annually
Employer
Company – Annapurna Labs (U.S.) Inc. – D63