Senior Performance Co-Design Engineer, TPU

Google Inc.

Sunnyvale (CA)

On-site

USD 174,000 - 252,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. The TPU Chip Architecture and Performance Co-design team optimizes Google's custom AI silicon for next-generation ML models, including LLMs, on our own hardware.

As a Senior Performance Co-design Engineer, you will focus on LLM Serving Studies, analyzing and optimizing serving performance on our custom hardware, and collaborating with hardware architects to

Qualifications

  • Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, a related field, or equivalent practical experience.
  • 5 years of experience in performance modeling/engineering, computer architecture, co-design, or systems engineering.
  • Experience programming in C++ or Python.

Responsibilities

  • Conduct comprehensive serving performance studies on current and emerging LLMs, including 1P/3P models.
  • Develop and maintain advanced simulation, profiling, and modeling tools to identify bottlenecks and project serving workload performance.
  • Partner with model researchers, software and hardware teams to co-design architectural improvements for LLM inference latency and throughput.
  • Drive data-backed decisions that influence the roadmap for future TPU/Cloud Silicon architectures.

Job description

Senior Performance Co-Design Engineer, LLM Serving
  • link Copy link

corporate_fare Google place Sunnyvale, CA, USA

Mid

Experience driving progress, solving problems, and mentoring more junior team members; deeper expertise and applied knowledge within relevant area.

  • Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, a related field, or equivalent practical experience.
  • 5 years of experience in performance modeling/engineering, computer architecture, co-design, or systems engineering.
  • Experience programming in C++ or Python.
Preferred qualifications:
  • Master's degree or PhD in Electrical Engineering, Computer Engineering or Computer Science, with an emphasis on computer architecture.
  • Experience with hardware/software co-design problems, especially performance analysis and identification at the pre-silicon stage.
  • Experience enabling and optimizing large-scale ML models (e.g., LLMs, large embedding models).
  • Experience with ML infrastructure, profiling tools, or deep learning inference/serving optimizations.
  • Familiarity with accelerator architectures.
About the job

Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action. In this role, you aren’t just building technology; you’re shaping the frontier of enterprise and driving the evolution of advanced models.

The TPU Chip Architecture and Performance Co-design team is at the forefront of optimizing Google's custom AI silicon for next-generation machine learning models.

As a Senior Performance Co-design Engineer, you will focus and conduct LLM Serving Studies. In this role, you will work on analyzing and optimizing the serving performance of emerging models and use cases on our custom hardware. You will also work closely with hardware architects to influence the evolution of Google’s custom ML accelerators.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We’re the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

About the job

Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action. In this role, you aren’t just building technology; you’re shaping the frontier of enterprise and driving the evolution of advanced models.

The TPU Chip Architecture and Performance Co-design team is at the forefront of optimizing Google's custom AI silicon for next-generation machine learning models.

As a Senior Performance Co-design Engineer, you will focus and conduct LLM Serving Studies. In this role, you will work on analyzing and optimizing the serving performance of emerging models and use cases on our custom hardware. You will also work closely with hardware architects to influence the evolution of Google’s custom ML accelerators.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We’re the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google .

  • Conduct comprehensive serving performance studies on (current and emerging) LLMs (1P/3P models).
  • Develop and maintain advanced simulation, profiling, and modeling tools to identify bottlenecks, understand key characteristics and project serving workload performance.
  • Partner with model researchers, software and hardware teams to co-design architectural improvements tailored to Large Language Model (LLM) inference latency and throughput.
  • Drive data-backed decisions that influence the roadmap for future TPU/Cloud Silicon architectures.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy , Know your rights: workplace discrimination is illegal , Belonging at Google , and How we hire .

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Performance Co-Design Engineer, TPU
Senior Performance Co-Design Engineer, TPU

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Senior Performance Co-design Engineer, LLM Training
Senior Performance Co-design Engineer, LLM Training

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Performance Co-Design Engineer, Google Cloud TPU
Performance Co-Design Engineer, Google Cloud TPU

Google • Sunnyvale (CA)

On-site
USD 192,000 - 278,000
Senior Staff Performance Codesign Engineer, TPU
Senior Staff Performance Codesign Engineer, TPU

Google • Sunnyvale (CA)

On-site
USD 240,000 - 333,000
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity grants
Senior Performance Co-Design Engineer, LLM Serving
Senior Performance Co-Design Engineer, LLM Serving

Socket.dev • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Senior Staff Performance Codesign Engineer, TPU
Senior Staff Performance Codesign Engineer, TPU

Google Inc. • Sunnyvale (CA)

On-site
USD 240,000 - 333,000
Bonus target
Equity
Benefits
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and disa
Senior Performance Co-design Engineer, LLM Training
Senior Performance Co-design Engineer, LLM Training

Socket.dev • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Google • United States

On-site
USD 207,000 - 300,000
Health, dental, vision, life, and/ or
401(k) with company match
PTO 20 days per year
+4