Senior ML Performance Engineer - Fleet TPU

Google

Sunnyvale (CA)

On-site

USD 174,000 - 252,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grants

Job summary

Google Sunnyvale, CA, USA is seeking a Senior Software Engineer for Fleet-level ML Performance to drive end-to-end performance analysis of key ML workloads on future TPU systems, using advanced simulation tools and cross‑functional collaboration across the ML stack.

The role emphasizes building scalable performance estimation methods in C++ and Python, balancing power, performance, and cost while addressing reliability and scheduling in data centers; you’ll influence TPU platform architecture

Qualifications

  • Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • 5 years of experience in systems architecture or computer architecture, power and performance trade-off analysis, or data center, cloud, infrastructure hardware optimization .
  • Experience with Reliability, Availability, and Serviceability (RAS) features, paradigms, or architecture.

Responsibilities

  • Perform fleet‑level performance analysis of key ML workloads (e.g., Gemini) on future TPU systems using advanced simulation tools to evaluate hardware/software trade‑offs and guide next‑generation chip architecture.
  • Partner with teams across the ML stack including model researchers, compiler developers, systems engineers, and TPU architects to analyze and optimize performance across the design space.
  • Design and implement a unified ML Accelerator Platform Performance Estimation Methodology using C++ and Python to enable scalable performance and TCO projections across Google.
  • Collaborate cross‑functionally with data center, hardware architecture, and framework teams to define key hardware and software requirements for future AI infrastructure.

Skills

Systems architecture
RAS features
Data center optimization
Cloud infrastructure
Performance analysis
Python
C++

Education

Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field
Master's degree or PhD in Electrical Engineering, Computer Engineering or Computer Science, with emphasis on computer architecture

Tools

Python
C++

Job description

Google Sunnyvale, CA, USA is seeking a Senior Software Engineer for Fleet-level ML Performance to drive end-to-end performance analysis of key ML workloads on future TPU systems, using advanced simulation tools and cross‑functional collaboration across the ML stack.

The role emphasizes building scalable performance estimation methods in C++ and Python, balancing power, performance, and cost while addressing reliability and scheduling in data centers; you’ll influence TPU platform architecture

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior TPU Performance Co-Design Engineer
Senior TPU Performance Co-Design Engineer

Google • Sunnyvale (CA)

On-site
USD 240,000 - 333,000
Staff ML Systems Co-Designer for TPU Performance
Staff ML Systems Co-Designer for TPU Performance

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and(dis)
401(k) with company match
Paid Time Off: 20 days/year
+4
ML Performance Engineering Manager
ML Performance Engineering Manager

Google • United States

On-site
USD 207,000 - 300,000
Health, dental, vision
401(k) with company match
Paid time off 20 days
+4
Senior ML Infra Engineer - TPU Health & Diagnostics
Senior ML Infra Engineer - TPU Health & Diagnostics

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Health insurance
Dental insurance
Vision insurance
+6
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Socket.dev • Sunnyvale (CA)

Hybrid
USD 174,000 - 252,000
Senior TPU Performance Architect for AI Systems
Senior TPU Performance Architect for AI Systems

Socket.dev • Sunnyvale (CA)

Hybrid
USD 174,000 - 252,000
Staff ML Systems Co-Design Engineer
Staff ML Systems Co-Design Engineer

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff TPU Performance Engineer: Optimize Large-Scale ML
Staff TPU Performance Engineer: Optimize Large-Scale ML

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
Dental, Vision, Life, Disability
401(k) with company match
+5
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity grants
Engineering Manager: ML Performance & TPU Optimizations
Engineering Manager: ML Performance & TPU Optimizations

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000