Senior AI Infra Platform Engineer: Cloud Hardware & Debug

Google

Sunnyvale (CA)

On-site

USD 188,000 - 274,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with company match
Paid time off (vacation) 20 days/year
Sick time 40 hours/year
Maternity leave 28-30 weeks
Baby bonding leave 18 weeks
13 paid holidays per year

Job summary

Google is hiring for an AI Infrastructure Platform Application Engineer (Hardware Engineer) to diagnose and resolve complex hardware observations, drive deep analysis and root-cause fixes across AI/ML infrastructure.

You will partner with engineering and product teams to continuously improve hardware platforms and software for customers using Google Cloud. The role requires extensive experience with server hardware, Linux, and debugging across the stack.

Qualifications

  • Bachelor's degree or equivalent practical experience in a technical field.
  • 6 years of experience with technical infrastructure deployment, maintenance, and troubleshooting.
  • 6 years of debug/validation experience with CPU, dGPU, or TPU.
  • 5 years of hardware debugging (silicon/platform IO/interface/memory).
  • Experience debugging across hardware/software stack (hardware faults, low-level software, networking, virtualization, kernel drivers, firmware, or performance).
  • Experience with Linux/Unix systems and debugging across hardware/software boundaries on enterprise-grade server infrastructure.

Responsibilities

  • Manage customer problems through effective diagnosis, resolution, or new tool development to improve AI/ML infrastructure productivity.
  • Collaborate with Product, Quality, and Engineering teams to improve products and support Site Reliability Engineering teams.
  • Debug platform hardware and silicon-related issues to drive root-cause resolution and permanent improvements.
  • Understand AI/ML workloads and hardware architectures; troubleshoot, reproduce, identify causes, and build faster diagnosis tools.
  • Act as a consultant for internal stakeholders to resolve deployment and operational challenges in AI infrastructure environments.

Skills

Platform debugging
Hardware debugging
AI/ML hardware
Linux/Unix
Cross-stack debugging

Education

Bachelor's degree in Computer Science or related field

Tools

Linux/Unix systems
CPU/dGPU/TPU hardware
Kubernetes/Slurm familiarity

Job description

Google is hiring for an AI Infrastructure Platform Application Engineer (Hardware Engineer) to diagnose and resolve complex hardware observations, drive deep analysis and root-cause fixes across AI/ML infrastructure.

You will partner with engineering and product teams to continuously improve hardware platforms and software for customers using Google Cloud. The role requires extensive experience with server hardware, Linux, and debugging across the stack.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Platform Engineer (Cloud&Hardware)
Senior AI Infrastructure Platform Engineer (Cloud&Hardware)

Google • Kirkland (WA)

On-site
USD 188,000 - 274,000
Health insurance
401(k) with company match
Paid time off
+4
AI Platform & Hardware Engineer
AI Platform & Hardware Engineer

Google • Austin (TX)

On-site
USD 159,000 - 230,000
Health insurance
Dental insurance
Vision insurance
+8
AI Cloud Infrastructure Platform Engineer
AI Cloud Infrastructure Platform Engineer

Google • Kirkland (WA)

Hybrid
USD 159,000 - 230,000
Health insurance
401(k) with company match
Paid time off
+4
Senior AI Infrastructure Platform Engineer
Senior AI Infrastructure Platform Engineer

Google LLC • Kirkland (WA), Sunnyvale (CA)

On-site
USD 164,000 - 274,000
Health, dental, vision, life, and long
401(k) with company match
PTO: 20 days/year
+4
Field AI Infra Engineer — Hardware & Cloud
Field AI Infra Engineer — Hardware & Cloud

Google • Kirkland (WA)

On-site
USD 132,000 - 189,000
Health/dental/vision insurance
401(k) with company match
Paid time off 20 days/year
+3
Lead AI Hardware Field Engineering Manager
Lead AI Hardware Field Engineering Manager

Google • Kirkland (WA)

On-site
USD 236,000 - 329,000
Health and welfare benefits
401(k) with company match
Paid time off 20 days
+4
Senior Platform Application Engineer, Cloud AI Infrastructure
Senior Platform Application Engineer, Cloud AI Infrastructure

Google • Sunnyvale (CA)

On-site
USD 188,000 - 274,000
Health insurance
401(k) with company match
Paid time off (vacation) 20 days/year
+4
AI Infrastructure Solutions Engineer
AI Infrastructure Solutions Engineer

Jobs in JS • Kirkland (WA)

On-site
USD 102,000 - 144,000
Equity
Benefits
Senior Platform Application Engineer, Cloud AI Infrastructure
Senior Platform Application Engineer, Cloud AI Infrastructure

Google • Kirkland (WA)

On-site
USD 188,000 - 274,000
Health insurance
401(k) with company match
Paid time off
+4
Platform Application Engineer, Cloud AI Infrastructure
Platform Application Engineer, Cloud AI Infrastructure

Google • Austin (TX)

On-site
USD 159,000 - 230,000
Health insurance
Dental insurance
Vision insurance
+8