Staff Software Engineer, TPU, Performance

Google

Ionia (NY)

On-site

USD 207,000 - 300,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale and extend well beyond web search.

We're looking for engineers who bring fresh ideas across AI, NLP, and large-scale systems; the roles span many domains and projects. Google's Core Machine Learning team focuses on TPUs, Gemini, and OSS ML models to optimize performance on

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience with ML design and ML infrastructure (model deployment, evaluation, data processing, debugging, fine tuning).
  • 5 years of experience testing and launching software products, and 3 years of experience with software design and architecture.

Responsibilities

  • Identify and maintain ML training and serving benchmarks that are representative to Google production and broader ML industry.
  • Achieve performance for customer launches, and in case of third-party/Open-Source Software models, for engaged benchmark submissions.
  • Use the benchmarks to identify performance opportunities and drive out-of-the-box performance toward improving the compiler, runtime, etc., in collaboration with those teams.
  • Engage with Google Product teams and researchers to solve their performance problems (e.g., onboard new ML models and products on Google new TPU hardware).
  • Analyze performance and efficiency metrics to identify bottlenecks, design, and implement solutions at Google fleet-wide scale.

Skills

Advanced software development
ML design & infrastructure
Speech/Audio/RL/ML acceleration
Testing & product launches

Education

Bachelor's degree or equivalent practical experience

Tools

CUDA/OpenCL
OpenXLA/MLIR/Triton
TensorFlow/PyTorch

Job description

Staff Software Engineer, TPU, Performance

Locations: Sunnyvale, CA, USA; New York, NY, USA

Level: Advanced

Advanced Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders; deep expertise in domain.

Important Note

In most instances, this position requires in-person interviews as part of the hiring process. Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Sunnyvale, CA, USA; New York, NY, USA.

Minimum Qualifications
  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience with one or more of the following: speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.
  • 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • 5 years of experience testing and launching software products, and 3 years of experience with software design and architecture.
Preferred Qualifications
  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • 8 years of experience with data structures and algorithms.
  • Experience with machine learning, compiler optimization, code generation, and runtime systems for GPU architectures (OpenXLA, MLIR, Triton, etc).
  • Experience in tailoring algorithms and ML models to exploit ML accelerator architecture strengths and minimize weaknesses.
  • Experience in low-level GPU programming (CUDA, OpenCL, etc.) and performance tuning techniques.
  • Understanding of modern Graphics Processing Unit (GPU), TPU or other ML accelerator architectures, memory hierarchies, and performance bottlenecks.
About the Job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google's needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

Google's Core Machine Learning (ML) organization is seeking software engineers to join the team known for pioneering work with Tensor Processing Units (TPUs). In this role, you will work on Gemini, as well as industry-leading open-source models, to understand model architecture and optimize the performance of these Machine Learning (ML) models on TPU systems for both Just After eXecution (JAX) and PyTorch platforms. You will improve the performance of ever-evolving ML workloads, achieving results. These fundamental efforts will influence next-generation (next-gen) TPU architectures via partnerships, ensuring performance for Gemini and Open-Source Software (OSS) Machine Learning (ML) models. The Core team builds the technical foundation behind Google's flagship products. We are owners and advocates for the underlying design elements, developer platforms, product components, and infrastructure at Google. These are the essential building blocks for excellent, safe, and coherent experiences for our users and drive the pace of innovation for every developer. We look across Google's products to build central solutions, break down technical barriers and strengthen existing systems. As the Core team, we have a mandate and a unique opportunity to impact important technical decisions across the company. Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US Salary : $207,000 - $300,000 (USD) + 20% bonus target + equity + benefits

Responsibilities
  • Identify and maintain ML training and serving benchmarks that are representative to Google production and broader ML industry.
  • Achieve performance for customer launches, and in case of third-party/Open-Source Software (3P/OSS) models, for engaged benchmark submissions (ML commons, InferenceMAX, etc.).
  • Use the benchmarks to identify performance opportunities and drive out-of-the-box performance toward improving the compiler, runtime, etc., in collaboration with those teams.
  • Engage with Google Product teams and researchers to solve their performance problems (e.g., onboard new ML models and products on Google new TPU hardware, enabling larger models (giant models) to train efficiently on a very large-scale (i.e., thousands of TPUs).
  • Analyze performance and efficiency metrics to identify bottlenecks, design, and implement solutions at Google fleet-wide scale.

Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy , Know your rights: workplace discrimination is illegal , Belonging at Google, and How we hire.

If you have a need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, TPU, Performance
Staff Software Engineer, TPU, Performance

Google • New York (NY)

On-site
USD 207,000 - 300,000
Staff Software Engineer, TPU, Performance
Staff Software Engineer, TPU, Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target 20%
Company benefits
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google Inc. • Mountain View (CA)

On-site
USD 207,000 - 300,000
Software Engineer III, TPU Performance, Hardware and Software Codesign
Software Engineer III, TPU Performance, Hardware and Software Codesign

Google Inc. • Sunnyvale (CA)

On-site
USD 147,000 - 210,000
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Google • United States

On-site
USD 174,000 - 252,000
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and(dis)
401(k) with company match
Paid Time Off: 20 days/year
+4
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid time off
+4
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, ML Frameworks
Staff Software Engineer, ML Frameworks

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Staff Software Engineer, Machine Learning, ML Training
Senior Staff Software Engineer, Machine Learning, ML Training

Google • Town of Montana (WI)

On-site
USD 262,000 - 364,000
Health/dental/vision
401(k) match
Paid time off
+4