Engineering Manager, GPU Reliability & Accelerators

Google Inc.

Sunnyvale, Northern (CA, KY)

Hybrid

USD 207,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Bonus target
Benefits

Job summary

Google Inc. in Sunnyvale, CA, is hiring an Engineering Manager to lead GPU Reliability for Accelerators. This role focuses on guiding a team of engineers, setting technical direction, and ensuring reliable performance of GPU systems across large-scale infrastructure.

You will oversee architecture, incident response, and cross-team collaboration to deliver scalable, reliable AI infrastructure, with in-depth experience in embedded systems and GPU software development.

Qualifications

  • Bachelor's degree or equivalent practical experience required.
  • 8 years of experience programming in C++, Java, Python, Kotlin or Go.
  • 5 years of experience with software architecture and embedded systems.
  • 3 years of experience in a technical leadership role.
  • 2 years of experience in a people management or team leadership role.
  • Experience with GPU programming, systems reliability and computer architecture.

Responsibilities

  • Architect the GPU Reliability Systems Operations and Tooling SW organization to transition from NPI-specific task forces into a centralized organization.
  • Drive the GPU support strategy by developing a roadmap for fleet health reliability, capacity turn-up, and automated health management.
  • Establish and enforce Service Level Objectives and Service Level Indicators for GPU Pod availability and performance.
  • Sponsor automation and toil reduction efforts by driving the development of advanced telemetry and debugging tooling.
  • Manage high-severity escalations for critical hardware and software issues, leading incident response and blameless post-mortems.

Skills

C++
Java
Python
Kotlin
Go
Software architecture
Embedded systems
Technical leadership
People management
GPU programming
System reliability
Computer architecture

Education

Bachelor's degree or equivalent practical experience
Master's degree or PhD in Computer Science or related field

Job description

Google Inc. in Sunnyvale, CA, is hiring an Engineering Manager to lead GPU Reliability for Accelerators. This role focuses on guiding a team of engineers, setting technical direction, and ensuring reliable performance of GPU systems across large-scale infrastructure.

You will oversee architecture, incident response, and cross-team collaboration to deliver scalable, reliable AI infrastructure, with in-depth experience in embedded systems and GPU software development.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Reliability Engineering Manager
GPU Reliability Engineering Manager

Google Inc. • Seattle (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) match
Paid time off
+4
Engineering Manager, GPU Reliability, Accelerators
Engineering Manager, GPU Reliability, Accelerators

Google Inc. • Sunnyvale (CA), Northern (KY)

Hybrid
USD 207,000 - 300,000
Equity
Bonus target
Benefits
Senior GPU System Software Engineer - AI & Infrastructure
Senior GPU System Software Engineer - AI & Infrastructure

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Senior AI/ML Software Architect — GPU Infrastructure
Senior AI/ML Software Architect — GPU Infrastructure

Google • Seattle (WA)

On-site
USD 262,000 - 364,000
Health insurance
Dental insurance
Vision insurance
+8
Technical Program Manager – AI/ML GPU Systems
Technical Program Manager – AI/ML GPU Systems

Google Inc. • Sunnyvale (CA)

On-site
USD 192,000 - 278,000
Health, dental, vision, life insurance
401(k) with company match
Paid time off 20 days/year
+4
Senior GPU System Software Engineer
Senior GPU System Software Engineer

Google • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Equity grants
Bonus target
Software Engineering Manager, GPU Reliability
Software Engineering Manager, GPU Reliability

Google Inc. • Seattle (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) match
Paid time off
+4
Lead GPU Architect for AI Accelerators & Clusters
Lead GPU Architect for AI Accelerators & Clusters

EngineersOfAI • Milpitas (CA)

On-site
USD 140,000 - 190,000
Tech Lead, Accelerator Platforms & HPC
Tech Lead, Accelerator Platforms & HPC

Google • Town of Montana (WI)

On-site
USD 262,000 - 364,000
Technical Program Manager, AI/ML GPU Systems
Technical Program Manager, AI/ML GPU Systems

Google • Kirkland (WA)

On-site
USD 192,000 - 278,000
Health insurance
Dental insurance
Vision insurance
+8