Principal/Distinguished Engineer, AI Capacity Delivery

Socket.dev

Sunnyvale (CA)

On-site

USD 364,000 - 505,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Bonus program
Benefits

Job summary

Google is seeking a Principal/ Distinguished Engineer for AI Capacity Delivery to bridge low-level hardware design with intelligent software automation. You will architect the intelligent control plane that manages and validates Google’s next-generation AI fleet, including TPUs and GPUs, delivering highly available, cost-optimized data center infrastructure at scale.

You will build adaptive ML-driven systems that analyze hardware behavior, predict anomalies, and configure hardware at scale.

Qualifications

  • 15 years of professional software engineering experience
  • Experience developing low-level systems software, virtualization layers, hypervisors, kernel extensions, or software that interfaces with bare-metal environments
  • Strong collaboration and cross-functional influence in delivering complex infrastructure projects

Responsibilities

  • Lead and architect the technical strategy for high-performance software control planes that orchestrate hardware and distributed data center environments.
  • Design and deploy self-training ML-driven monitoring and telemetry systems that automate system validation.
  • Write high-efficiency systems software to automate fleet initialization and production readiness.
  • Influence long-term hardware-software co-design roadmaps with TPU, silicon, and distributed systems teams.

Education

Bachelor’s degree in Computer Science, Computer Engineering, or related field
Master’s degree or PhD in Computer Science, Computer Engineering, or related field

Tools

Hypervisors
Kernel extensions

Job description

Minimum qualifications:
  • Bachelor’s degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience.
  • 15 years of professional software engineering experience, with a focus on systems programming, distributed systems, or high-performance infrastructure orchestration.
  • Experience developing low-level systems software, virtualization layers, hypervisors, kernel extensions, or software that interfaces directly with physical infrastructure/bare-metal environments.
Preferred qualifications:
  • Master’s degree or PhD in Computer Science, Computer Engineering, or related field.
  • Significant background in machine learning platforms and infrastructure, with experience applying AI/ML models to automated system testing, hardware qualification, or predictive system reliability.
  • Technical expertise building and validating large-scale distributed systems, bare-metal environments, or hyperscale data center networks.
  • Exceptional collaborator with a proven track record of influencing cross-functional teams to drive delivery.
  • Track record at a hyperscale cloud provider or AI chipmaker, working on low-level infrastructure software (e.g., hypervisors, hardware abstraction layers, or custom kernels).
About the job:

As a Principal/Distinguished Engineer for AI Capacity Delivery, you will bridge the gap between low-level hardware design and intelligent software automation. You will architect the intelligent software control plane that manages, validates, and initializes Google’s next-generation global AI fleet. Your core mission is to drastically accelerate the onboarding and validation of Google’s next-generation AI fleet, including TPUs and GPUs. You will deliver and operate highly available, cost optimized data center infrastructure at speed and scale.

You will leverage modern machine learning platforms to build adaptive, self-training systems that analyze physical fleet behavior, predict anomalies, and automatically configure and validate hardware at scale. This person must develop the necessary hardware qualification tests that ensure that our fleet is reliable, properly configured and healthy enough to perform its mission. You will leverage advanced ML platforms to develop self-training systems that analyze physical hardware behavior and guide engineering teams to write highly optimized software for testing.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We’re the driving force behind Google’s groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $364000 - $505000 (USD) + 40% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities:
  • Lead and architecture the technical strategy for high-performance software control planes that orchestrate, validate, and manage bare-metal systems, custom AI silicon, and distributed data center environments.
  • Design and deploy self-training ML-driven monitoring and telemetry systems that automate system validation, replacing legacy test suites with adaptive, self-correcting frameworks.
  • Write high-efficiency systems software that automates fleet initialization, minimizing the time it takes to transition physical hardware into live, reliable production capacity.
  • Influence the long-term technical roadmap for hardware-software co-design, collaborating closely with TPU, silicon, and distributed systems engineering teams to ensure the reliability and scale of Google’s planet scale infrastructure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal/Distinguished Engineer, AI Capacity Delivery
Principal/Distinguished Engineer, AI Capacity Delivery

Google • Sunnyvale (CA)

On-site
USD 364,000 - 505,000
Principal/Distinguished Engineer, AI Capacity Delivery
Principal/Distinguished Engineer, AI Capacity Delivery

Google Inc. • Sunnyvale (CA)

On-site
USD 364,000 - 505,000
Tech Lead, TPU AI Infrastructure
Tech Lead, TPU AI Infrastructure

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Software Engineer, AI/TPU Infrastructure Firmware, On-prem
Senior Software Engineer, AI/TPU Infrastructure Firmware, On-prem

Socket.dev • Kirkland (WA)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Tech Lead, TPU AI Infrastructure
Tech Lead, TPU AI Infrastructure

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Bonus target
Equity
Benefits
Principal Engineer, AI Infra & Bare-Metal Orchestration
Principal Engineer, AI Infra & Bare-Metal Orchestration

Socket.dev • Sunnyvale (CA)

On-site
USD 364,000 - 505,000
Equity
Bonus program
Benefits
Senior Hardware Engineering Manager, Emergent AI Infrastructure
Senior Hardware Engineering Manager, Emergent AI Infrastructure

Socket.dev • Sunnyvale (CA)

On-site
USD 236,000 - 329,000
Principal AI Systems Engineer, Hyper-scale Infra
Principal AI Systems Engineer, Hyper-scale Infra

Google Inc. • Sunnyvale (CA)

On-site
USD 364,000 - 505,000
Tech Lead, TPU AI Infrastructure
Tech Lead, TPU AI Infrastructure

Google Inc. • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Benefits
Staff Software Engineer, AI/ML Infrastructure, TPU Supercomputers
Staff Software Engineer, AI/ML Infrastructure, TPU Supercomputers

Google • Sunnyvale (CA)

On-site
USD 207,000 - 301,000
20% bonus target
Equity
Comprehensive benefits package