Staff TPU Performance Engineer: Optimize Large-Scale ML

Google

Kirkland (WA)

On-site

USD 207,000 - 300,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental, Vision, Life, Disability
401(k) with company match
Paid Time Off 20 days per year
Sick Time 40 hours/year (Seattle 69)
Maternity Leave 28-30 weeks
Baby Bonding Leave 18 weeks
Holidays 13 days

Job summary

Google is seeking software engineers to advance next-generation technologies that scale to billions of users. You will work on projects across JAX and PyTorch, focusing on TPU fleet efficiency and performance optimization, collaborating with product teams and researchers to solve complex ML deployment challenges.

The role emphasizes designing, developing, testing, deploying, and maintaining software solutions with opportunities to switch teams as Google grows.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience testing, and launching software products.
  • 5 years of experience with performance, large-scale systems data analysis, visualization tools, or debugging.
  • 3 years of experience with software design and architecture.
  • Experience with ML performance analysis and benchmarking.

Responsibilities

  • Learn more about benefits at Google .
  • Focus on TPU fleet efficiency analysis and performance optimization, while identifying and maintaining ML training and serving benchmarks.
  • Use benchmarks to identify performance opportunities and drive out-of-the-box performance by improving the compiler, runtime, etc. in collaboration with partner teams.
  • Collaborate with Google product teams and researchers to solve performance problems, such as onboarding new ML models and products onto new TPU hardware to enable larger models to train efficiently at a very large scale.
  • Analyze performance and efficiency metrics to identify bottlenecks, design, and implement solutions at Google fleet-wide scale.
  • Explore model and data efficiency techniques i.e., model co-design, quantization, and sparsity.

Skills

ML performance analysis
Performance analysis
Debugging

Education

Bachelor's degree or equivalent practical experience
Master’s degree or PhD in Engineering, CS, or related field

Tools

OpenXLA
NVIDIA/AMD architectures

Job description

Google is seeking software engineers to advance next-generation technologies that scale to billions of users. You will work on projects across JAX and PyTorch, focusing on TPU fleet efficiency and performance optimization, collaborating with product teams and researchers to solve complex ML deployment challenges.

The role emphasizes designing, developing, testing, deploying, and maintaining software solutions with opportunities to switch teams as Google grows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff TPU Performance Architect for Large-Scale ML
Staff TPU Performance Architect for Large-Scale ML

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Staff Software Engineer, TPU Performance
Senior Staff Software Engineer, TPU Performance

Socket.dev • Sunnyvale (CA)

On-site
USD 262,000 - 365,000
Equity
Benefits
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Engineering Manager: ML Performance & TPU Optimizations
Engineering Manager: ML Performance & TPU Optimizations

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior ML Systems Architect – TPU & AI Infra
Senior ML Systems Architect – TPU & AI Infra

Socket.dev • Sunnyvale (CA)

On-site
USD 262,000 - 365,000
Equity
Benefits
Staff ML Systems Co-Design Engineer
Staff ML Systems Co-Design Engineer

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Engineering Manager, ML Performance
Engineering Manager, ML Performance

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior TPU Performance Architect for AI Systems
Senior TPU Performance Architect for AI Systems

Socket.dev • Sunnyvale (CA)

Hybrid
USD 174,000 - 252,000
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Socket.dev • Sunnyvale (CA)

Hybrid
USD 174,000 - 252,000
Staff AI Engines Engineer: TPU Inference & ML Runtime
Staff AI Engines Engineer: TPU Inference & ML Runtime

Socket.dev • Mountain View (CA)

On-site
USD 207,000 - 300,000