Principal/Distinguished Engineer, AI Capacity Delivery

Google Inc.

Sunnyvale (CA)

On-site

USD 364,000 - 505,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Google Inc. is seeking a Principal/Distinguished Engineer for AI Capacity Delivery to bridge hardware design with intelligent software automation.

You will architect a scalable control plane to manage, validate, and initialize Google’s AI fleet, including TPUs and GPUs, delivering high availability and cost efficiency. Lead the technical strategy for bare-metal systems, write efficient system software for fleet initialization, and collaborate with TPU/silicon teams to ensure reliability at

Qualifications

  • Bachelor’s degree or equivalent in CS/CE or related field.
  • 15 years of professional software engineering experience.
  • Experience with low-level systems software or virtualization/bare-metal environments.

Responsibilities

  • Design and deploy ML-driven monitoring and telemetry for validation.
  • Automate fleet initialization and rapid bring-up of hardware.
  • Influence long-term hardware-software co-design roadmaps with TPU, silicon, and distributed-systems teams.

Skills

Systems programming
Distributed systems
Low-level software
Kernel interfaces

Education

Bachelor's degree required
Master's or PhD preferred

Job description

Principal/Distinguished Engineer, AI Capacity Delivery
  • Bachelor’s degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience.
  • 15 years of professional software engineering experience, with a focus on systems programming, distributed systems, or high-performance infrastructure orchestration.
  • Experience developing low-level systems software, virtualization layers, hypervisors, kernel extensions, or software that interfaces directly with physical infrastructure/bare-metal environments.
Preferred qualifications:
  • Master’s degree or PhD in Computer Science, Computer Engineering, or related field.
  • Significant background in machine learning platforms and infrastructure, with experience applying AI/ML models to automated system testing, hardware qualification, or predictive system reliability.
  • Technical expertise building and validating large-scale distributed systems, bare-metal environments, or hyperscale data center networks.
  • Exceptional collaborator with a proven track record of influencing cross-functional teams to drive delivery.
  • Track record at a hyperscale cloud provider or AI chipmaker, working on low-level infrastructure software (e.g., hypervisors, hardware abstraction layers, or custom kernels).
About the job

As a Principal/Distinguished Engineer for AI Capacity Delivery, you will bridge the gap between low-level hardware design and intelligent software automation. You will architect the intelligent software control plane that manages, validates, and initializes Google’s next-generation global AI fleet. Your core mission is to drastically accelerate the onboarding and validation of Google’s next-generation AI fleet, including TPUs and GPUs. You will deliver and operate highly available, cost optimized data center infrastructure at speed and scale.

You will leverage modern machine learning platforms to build adaptive, self-training systems that analyze physical fleet behavior, predict anomalies, and automatically configure and validate hardware at scale. This person must develop the necessary hardware qualification tests that ensure that our fleet is reliable, properly configured and healthy enough to perform its mission. You will leverage advanced ML platforms to develop self-training systems that analyze physical hardware behavior and guide engineering teams to write highly optimized software for testing.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $364000 - $505000 (USD) + 40% bonus target + equity + benefits. Learn more about benefits at Google.

Lead and architecture the technical strategy for high-performance software control planes that orchestrate, validate, and manage bare-metal systems, custom AI silicon, and distributed data center environments.

  • Design and deploy self-training ML-driven monitoring and telemetry systems that automate system validation, replacing legacy test suites with adaptive, self-correcting frameworks.
  • Write high-efficiency systems software that automates fleet initialization, minimizing the time it takes to transition physical hardware into live, reliable production capacity.
  • Influence the long-term technical roadmap for hardware-software co-design, collaborating closely with TPU, silicon, and distributed systems engineering teams to ensure the reliability and scale of Google's planet scale infrastructure.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy, Know your rights: workplace discrimination is illegal, Belonging at Google, and How we hire.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal/Distinguished Engineer, AI Capacity Delivery
Principal/Distinguished Engineer, AI Capacity Delivery

Google • Sunnyvale (CA)

On-site
USD 364,000 - 505,000
Tech Lead, TPU AI Infrastructure
Tech Lead, TPU AI Infrastructure

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Benefits
Senior Software Engineer, ML Infrastructure, Control Plane
Senior Software Engineer, ML Infrastructure, Control Plane

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Tech Lead Manager, Staff Software Engineering, XProf
Tech Lead Manager, Staff Software Engineering, XProf

Google • Sunnyvale (CA)

On-site
USD 207,000 - 301,000
Staff Software Engineer, Infrastructure, Google Cloud AI
Staff Software Engineer, Infrastructure, Google Cloud AI

Google • United States

On-site
USD 207,000 - 300,000
Staff Software Engineer, Network Health
Staff Software Engineer, Network Health

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 301,000
20% Bonus Target
Equity Options
Comprehensive Benefits
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Google • United States

On-site
USD 174,000 - 252,000
Technical Lead Manager, PIE Core Systems and Capacity
Technical Lead Manager, PIE Core Systems and Capacity

Google • Austin (TX)

On-site
USD 207,000 - 301,000
Staff Software Engineer, Data Center Network Switches
Staff Software Engineer, Data Center Network Switches

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, Data Center Resource Modeling
Staff Software Engineer, Data Center Resource Modeling

Google Inc. • Sunnyvale (CA)

On-site
USD 210,000 - 300,000