GPU Reliability Engineering Manager

Google Inc.

Seattle (WA)

On-site

USD 207,000 - 300,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) match
Paid time off
Sick time
Maternity leave
Baby bonding leave
Paid holidays

Job summary

Google Seattle, WA, USA is seeking an Software Engineering Manager, GPU Reliability to lead a team of engineers responsible for the reliability and performance of Google's GPU infrastructure and AI platforms. You will own technical roadmaps, mentor engineers, and influence product strategy across multiple teams.

The role requires advanced leadership, 8+ years in software development, and 2+ years in people management, with hands-on experience in ML design and infrastructure.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience leading technical project strategy, ML design, and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.
  • 3 years of experience in a technical leadership role.
  • 2 years of experience in a people management, supervision/team leadership role.

Responsibilities

  • Lead, mentor, and grow a team of software engineers, fostering a collaborative culture and supporting career development.
  • Execute technical roadmaps for the GPU ecosystem, anticipating market shifts to advance Google Cloud's AI infrastructure.
  • Oversee the accelerator solution lifecycle, driving root-cause analysis and resolving production issues to guarantee fleet reliability.
  • Collaborate with customers as a technical advocate to resolve critical challenges and translate feedback into platform enhancements.
  • Author and review technical specifications to maintain architectural consistency, high code quality, and long-term platform maintainability.

Skills

Outcome ownership
Leadership
ML design
ML infrastructure
Data processing
Debugging
Fine-tuning
Software testing
System architecture
People management

Education

Bachelor's degree or equivalent

Job description

Google Seattle, WA, USA is seeking an Software Engineering Manager, GPU Reliability to lead a team of engineers responsible for the reliability and performance of Google's GPU infrastructure and AI platforms. You will own technical roadmaps, mentor engineers, and influence product strategy across multiple teams.

The role requires advanced leadership, 8+ years in software development, and 2+ years in people management, with hands-on experience in ML design and infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Reliability Engineering Manager – AI Platform Leader
GPU Reliability Engineering Manager – AI Platform Leader

Google • Seattle (WA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and long
Disability insurance
401(k) with company match
+2
Software Engineering Manager, GPU Reliability
Software Engineering Manager, GPU Reliability

Google • Seattle (WA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and long
Disability insurance
401(k) with company match
+2
Software Engineering Manager, GPU Reliability
Software Engineering Manager, GPU Reliability

Google Inc. • Seattle (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) match
Paid time off
+4
Senior AI/ML Software Architect — GPU Infrastructure
Senior AI/ML Software Architect — GPU Infrastructure

Google • Seattle (WA)

On-site
USD 262,000 - 364,000
Health insurance
Dental insurance
Vision insurance
+8
GPU Infrastructure TPM III - Drive Reliable Cloud AI at Scale
GPU Infrastructure TPM III - Drive Reliable Cloud AI at Scale

Epic Games (Portuguese) • Seattle (WA)

On-site
USD 156,000 - 229,000
Lead ML Performance Engineering Manager
Lead ML Performance Engineering Manager

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid Time Off: 20 days per year
+4
Senior GPU System Software Engineer
Senior GPU System Software Engineer

Google • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Equity grants
Bonus target
Engineering Manager, ML Efficiency & AI Systems
Engineering Manager, ML Efficiency & AI Systems

Socket.dev • Mountain View (CA)

On-site
USD 207,000 - 300,000
Senior Software Engineer & SOC Architect - GPU/ML
Senior Software Engineer & SOC Architect - GPU/ML

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Engineering Manager, ML Performance
Engineering Manager, ML Performance

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid Time Off: 20 days per year
+4