Software Development Engineer, EC2 UltraServer Availability (AWS)

Amazon

Seattle (WA)

On-site

USD 143,700 - 194,400

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon in Seattle, WA seeks a Software Development Engineer II to design and maintain cloud-based repair and recovery workflows for NVIDIA UltraServer GPUs. The role collaborates with Capacity Management, Hardware Engineering, and Datacenter Operations to manage AI/ML infrastructure and scale GPU clusters.

Responsibilities include building AWS-based infrastructure, writing robust code with tests and reviews, and coordinating with downstream teams to troubleshoot failures while striving for

Qualifications

  • 3+ years of professional software development experience.
  • 2+ years of design or architecture experience on existing or new systems.
  • Experience programming with at least one software language.

Responsibilities

  • Collaborate with cross-functional teams to design stable, scalable infrastructure solutions.
  • Build cloud-based solutions using AWS services for scalable infrastructure frameworks.
  • Write high-quality, testable code with proper reviews and testing.
  • Develop and maintain repair and recovery workflows for UltraServer hosts.

Skills

Software development
System design
Programming language

Education

Bachelor's degree

Job description

Software Development Engineer, EC2 UltraServer Availability

Job ID: 10388616 | Amazon Development Center U.S., Inc.

The Software Development Engineer II will design, build, and maintain cloud-based repair and recovery workflows for NVIDIA GB200 / GB300 UltraServers. The role requires expertise in AWS services, system architecture, and cross‑functional collaboration with Capacity Management, Hardware Engineering, and Datacenter Operations to manage AI/ML infrastructure.

Key Responsibilities
  • By design and architecture, collaborate with cross‑functional teams to develop stable, logical, testable, and efficient infrastructure solutions.
  • Build cloud‑based solutions using AWS native services for scaling infrastructure frameworks.
  • Write high‑quality, maintainable code with proper testing and code reviews.
  • Develop and maintain the repair and recovery workflows for GB200 and GB300 UltraServer hosts.
  • Implement automation for diagnostic triage, hardware testing, cable validation, and testing processes.
  • Create observable systems with appropriate metrics and alarming.
  • Execute and monitor UltraServer workflows for repair and coordinate with downstream teams when troubleshooting failures.
  • Focus on operational excellence by identifying problems and proposing solutions.
  • Work with hardware and software integrations specific to GPU clusters and AI/ML training systems.
  • Manage network partition configurations for multi‑node GPU clusters.
  • Handle firmware validation and consistency checks across asset groups.
  • Collaborate with customers and stakeholders to convert business needs into technical designs.
  • Participate in code reviews and technical assessments.
A Day in the Life

This hands‑on position involves owning everything from requirements gathering, design, implementation, code reviews, incremental feature launches, operations, mentoring, to continuous improvement.

Basic Qualifications
  • 3+ years of professional software development experience.
  • 2+ years of design or architecture experience of new or existing systems.
  • Experience programming with at least one software programming language.
Preferred Qualifications
  • 3+ years of full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.
  • Bachelor’s degree in computer science or equivalent.
  • Knowledge of professional software engineering best practices for full software development life cycle.
Salary and Benefits

USA, WA, Seattle – 143,700.00 – 194,400.00 USD annually. The base salary range is listed below. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance, 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

Equal Employment Opportunity

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Development Engineer, EC2 UltraServer Delivery Team (AWS)
Software Development Engineer, EC2 UltraServer Delivery Team (AWS)

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
Sign-on payments
Restricted stock units (RSUs)
+3
Software Development Engineer, EC2 Trainium AI Infra
Software Development Engineer, EC2 Trainium AI Infra

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+2
Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
Senior Software Development Engineer, EC2 Trainium AI Infra
Senior Software Development Engineer, EC2 Trainium AI Infra

Amazon • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
401(k) matching
Paid time off
+1
SDE II — EC2 UltraServer Availability & AI/ML Infra
SDE II — EC2 UltraServer Availability & AI/ML Infra

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
401(k) matching
Paid time off
+1
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Generative AI & ML Servers
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Generative AI & ML Servers
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
RSUs
Health insurance
401(k) matching
Senior Hardware Development Manager, AWS Accelerator Servers
Senior Hardware Development Manager, AWS Accelerator Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 240,000 - 324,000
Health insurance
RSUs
401(k) matching
+2
Systems Development Engineer, AWS Generative AI & ML Servers
Systems Development Engineer, AWS Generative AI & ML Servers

Socket.dev • Austin (TX)

On-site
USD 129,000 - 175,000