Sr. Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering

Amazon Web Services (AWS)

Austin (TX)

On-site

USD 151,200 - 204,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon’s Annapurna Labs (U.S.) Inc. in Austin, TX is seeking engineers to build the backbone for Generative AI cloud.

You will work with cross‑functional teams to deliver next‑generation AWS platforms, owning complex server systems and driving reliability and scale. The role requires 4+ years in programming (C++, C#, Java, Python, Go, PowerShell, or Ruby) and 3+ years designing scalable systems, with experience across x86, ARM, and GPU/FPGA devices.

Qualifications

  • 4+ years of programming in at least one modern language (C++, C#, Java, Python, Go, PowerShell, Ruby).
  • 3+ years of professional software development experience.
  • 3+ years of designing or architecting systems for reliability and scaling.
  • Experience with x86, ARM and GPU/FPGA devices.
  • Knowledge of storage, networking, memory, and interface standards (I2C, IPMI, SPI, PCIe).
  • BS degree in CS/CE or equivalent work experience.

Responsibilities

  • Solve complex architectural problems across AWS platforms.
  • Own and scale team systems; write code to prevent customer impact.
  • Decompose server system testability, reliability, and diagnostics issues.
  • Diagnose and improve systems using hardware, software, and accelerator tech.
  • Create automation and AI-driven tools to support workflows.
  • Collaborate with SDEs, SDETs, TPMs, and managers across AWS.

Skills

C++
C#
Java
Python
Golang
PowerShell
Ruby

Education

BS in Computer Science or Computer Engineering or related field

Tools

IPMI
PCIe
I2C
SPI

Job description

Job Overview

Do you want to build the backbone of Generative AI cloud at AWS? Do you want to build the future of the cloud for AI training and inference, delivering continuous price‑performance improvements for multi‑billion‑variable LLMs? Join us in designing, delivering, and operating AWS cloud offerings that enable high‑performance, scalable AI/ML and HPC workloads.

Key Responsibilities

You will work with engineers across the company to deliver the next‑generation AWS platforms. Your responsibilities include:

  • Solving complex architectural problems that may not be defined before hand.
  • Owning the team’s systems, proactively identifying deficiencies, writing tactical code to solve issues before they impact customers, and scaling solutions.
  • Decomposing large server system testability, reliability, and diagnostics problems into manageable tasks or features, then leading delivery with other team members.
  • Utilizing hardware, software, system design, x86 (and ARM), GPU/FPGA knowledge, and modern storage, networking, and memory technologies to diagnose and improve systems.
  • Creating automation via agentic workflows, developing AI‑driven tools and workflows, and contributing to AI transformation.
  • Collaborating with SDEs, SDETs, TPMs, managers, and principals across AWS to drive high quality and reliability into future designs and accelerator server solutions.
Basic Qualifications
  • 4+ years of programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, or Ruby.
  • 3+ years of professional software development experience (non‑internship).
  • 3+ years of designing or architecting new and existing systems, focusing on reliability and scaling.
  • Experience with x86 architecture, as well as ARM and GPU/FPGA devices.
  • Knowledge of modern technology devices in storage, network, memory, and interface standards (I2C, IPMI, SPI, PCIe).
  • BS degree in Computer Science, Computer Engineering, or related technical degree, or equivalent work experience.
Preferred Qualifications
  • 7+ years or more of experience in software development, systems development, SRE, or resilience engineering.
  • 7+ years of SysDE or equivalent experience.
  • 7+ years of server systems debugging experience, including root cause analysis of complex server platforms.
  • Experience in improving durability, security, availability, and scalability of systems through exploration, diagnosis, and remediation.
  • Linux kernel and user‑space driver experience for PCIe and external devices.
  • System thinking: diagnosing interactions between discrete components of a server system and driving product improvements.
  • Strong focus on reliability, scale, and diagnostics; developing tactical and strategic tools using Python, Go, or C/C++.
  • Solid understanding of OS internals, including network and storage subsystems.
  • Master’s degree in Electrical Engineering, Computer Engineering, or related field.
  • Experience validating hardware, software, firmware, and drivers, and implementing test plans.
  • Experience with server validation, testing, root‑cause analysis, and coverage analysis.
  • Excellent diagnostics tools development experience with Python, Go, or C/C++ in a fast‑paced environment.
  • Extensive Linux knowledge.
Equal Employment Opportunity Statement

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Location and Compensation

USA, CA, Cupertino – 173,900.00 – 235,200.00 USD annually
USA, TX, Austin – 151,200.00 – 204,600.00 USD annually
USA, WA, Seattle – 151,200.00 – 204,600.00 USD annually

Employer

Company – Annapurna Labs (U.S.) Inc. – D63

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering
Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering

Amazon • Austin (TX)

On-site
USD 129,000 - 175,000
RSUs
Health insurance
Paid time off
AWS Systems Engineer: Generative AI & ML Servers
AWS Systems Engineer: Generative AI & ML Servers

Amazon • Austin (TX)

On-site
Sr. Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering
Sr. Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering

Amazon • Austin (TX)

On-site
USD 151,200 - 204,600
Comprehensive health insurance
401(k) matching
Paid time off
+1
Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering
Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Sr Hardware Development Engineer, High Performance AI & ML Servers
Sr Hardware Development Engineer, High Performance AI & ML Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 216,000
Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering
Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering

Amazon • Cupertino (CA)

On-site
USD 149,000 - 201,000
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 99,000 - 160,000
Health insurance
401(k) matching
Paid time off
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 123,000 - 185,000
Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML
Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 173,000 - 236,000
Sign-on bonuses and RSUs
Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML
Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 151,000 - 205,000