TPU Performance Engineer — ML Efficiency & Scale

Google

Sunnyvale (CA)

On-site

USD 207,000 - 300,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google is seeking software engineers for the AI and Infrastructure teams to push performance across large-scale systems and accelerator hardware. You will work across JAX and PyTorch to optimize production and research workloads, including Gemini models, across TPU fleets.

Responsibilities include designing and delivering high-performance software, collaborating with cross-functional teams, and driving scale, efficiency, and reliability in a fast-paced environment with opportunities to switch

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience testing and launching software products.
  • 5 years of experience with performance, large-scale systems data analysis, visualization tools, or debugging.
  • 3 years of experience with software design and architecture.
  • Experience with ML performance analysis and benchmarking.

Responsibilities

  • Focus on TPU fleet efficiency analysis and performance optimization, while identifying and maintaining ML training and serving benchmarks.
  • Use benchmarks to identify performance opportunities and drive out-of-the-box performance by improving the compiler, runtime, etc.
  • Collaborate with Google product teams and researchers to solve performance problems and onboard new ML models onto TPU hardware.
  • Analyze performance and efficiency metrics to identify bottlenecks and implement solutions at fleet-wide scale.
  • Explore model and data efficiency techniques such as model co-design, quantization, and sparsity.

Skills

Software development
Testing and launching software
Performance analysis
Large-scale data analysis
Visualization tools
Debugging
Software design
Software architecture
ML performance benchmarking

Education

Bachelor's degree or equivalent practical experience

Tools

OpenXLA

Job description

Google is seeking software engineers for the AI and Infrastructure teams to push performance across large-scale systems and accelerator hardware. You will work across JAX and PyTorch to optimize production and research workloads, including Gemini models, across TPU fleets.

Responsibilities include designing and delivering high-performance software, collaborating with cross-functional teams, and driving scale, efficiency, and reliability in a fast-paced environment with opportunities to switch

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff TPU Performance Engineer: Optimize Large-Scale ML
Staff TPU Performance Engineer: Optimize Large-Scale ML

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
Dental, Vision, Life, Disability
401(k) with company match
+5
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior TPU ML Performance Engineer
Senior TPU ML Performance Engineer

Google • United States

On-site
USD 174,000 - 252,000
Senior AI/ML Infrastructure Engineer for TPU Health
Senior AI/ML Infrastructure Engineer for TPU Health

Google • United States

On-site
USD 174,000 - 252,000
Health insurance
Dental insurance
Vision insurance
+8
TPU Kernel Engineer for High-Performance ML Systems
TPU Kernel Engineer for High-Performance ML Systems

SignalAI • New York (NY)

Hybrid
USD 280,000 - 850,000
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
Dental, Vision, Life, Disability
401(k) with company match
+5
Senior ML Performance Engineer - Fleet TPU
Senior ML Performance Engineer - Fleet TPU

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity grants
Senior AI/ML Infrastructure Engineer (TPU Health)
Senior AI/ML Infrastructure Engineer (TPU Health)

Google Inc. • Kirkland (WA), Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Health insurance
Retirement plan
Paid time off
+3
Senior AI/ML Infra Engineer — TPU Health Systems
Senior AI/ML Infra Engineer — TPU Health Systems

Google • Kirkland (WA)

On-site
USD 174,000 - 252,000
Health insurance
Retirement benefits
Paid time off (PTO)
+4
Senior TPU Performance Co-Design Engineer (LLM Serving)
Senior TPU Performance Co-Design Engineer (LLM Serving)

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000