Global Capacity Manager – TPU Focus

Jobtailor

California (MO)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Jobtailor is seeking a senior TPU capacity engineer to lead TPU pod fleets, design capacity management, and optimize multi-cloud environments for cost and performance.

The role requires deep Kubernetes and GCP TPU expertise, Go or Python production experience, and strong collaboration across teams to secure enterprise TPU capacity.

Qualifications

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
  • 5+ years of professional work experience in a high-growth environment, preferably at a hyperscaler (GCP, AWS, Azure) or a specialized accelerator provider.
  • Hands-on experience with Google Cloud TPUs, including pod slicing, ICI topology, JAX/XLA, and TPU-specific scheduling and fault handling
  • Deep expertise in Kubernetes, including taints, cordons, node draining, and custom operators
  • Demonstrated experience with Go or Python in a production-level environment
  • Strong financial literacy and ability to model complex trade-offs between capacity reliability and cost
  • High tenacity and collaborative mindset
  • Experience with additional non-NVIDIA accelerators, such as AWS Trainium/Inferentia or AMD Instinct, is nice to have
  • Familiarity with multi-accelerator scheduling and cost/performance tradeoff modeling is nice to have
  • Prior experience partnering with model performance or ML systems teams is nice to have

Responsibilities

  • Lead TPU pod fleets through acquisition, allocation, and maintenance
  • Execute complex workload migrations and sticky deployment drains across TPU topologies
  • Ensure deployment scheduling rules meet regional and compliance requirements
  • Design and implement the next version of Baseten's capacity management system
  • Scale TPU capacity alongside the existing GPU fleet
  • Build ROI models comparing TPU, GPU, and other accelerator options
  • Partner with MP, SRE, Infra, and FDE teams to tune workloads for TPU execution
  • Lead capacity-crunch incident response by reallocating and re-coordinating TPU workloads
  • Architect infrastructure readiness and deployment strategies for TPU clusters, including pod slicing and topology planning
  • Build multi-cloud capacity management systems to optimize cost and latency
  • Develop automated operators to identify, cordon, and repair unhealthy TPU pods
  • Partner with leadership and Google Cloud to secure dedicated TPU capacity for enterprise customers

Skills

Google Cloud TPU Management
Kubernetes Expertise
Go or Python Programming
Financial Modeling
Collaborative Mindset
Experience with Non-NVIDIA Accelerors
Multi-Accelerator Scheduling
Model Performance & ML Systems Liaison
TPU Pod Slicing
ICI Topology
JAX/XLA
Custom Operators
Node Draining

Education

Bachelor's/Master's/Ph.D. in Computer Science, Engineering, Mathematics, or related field

Tools

Google Cloud
AWS
Azure
Multi-Cloud Systems

Job description

  • Lead TPU pod fleets through acquisition, allocation, and maintenance
  • Execute complex workload migrations and sticky deployment drains across TPU topologies
  • Ensure deployment scheduling rules meet regional and compliance requirements
  • Design and implement the next version of Baseten's capacity management system
  • Scale TPU capacity alongside the existing GPU fleet
  • Build ROI models comparing TPU, GPU, and other accelerator options
  • Partner with MP, SRE, Infra, and FDE teams to tune workloads for TPU execution
  • Lead capacity-crunch incident response by reallocating and re-coordinating TPU workloads
  • Architect infrastructure readiness and deployment strategies for TPU clusters, including pod slicing and topology planning
  • Build multi-cloud capacity management systems to optimize cost and latency
  • Develop automated operators to identify, cordon, and repair unhealthy TPU pods
  • Partner with leadership and Google Cloud to secure dedicated TPU capacity for enterprise customers
Requirements
  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field
  • 5+ years of professional work experience in a high-growth environment, preferably at a hyperscaler (GCP, AWS, Azure) or a specialized accelerator provider
  • Hands-on experience with Google Cloud TPUs, including pod slicing, ICI topology, JAX/XLA, and TPU-specific scheduling and fault handling
  • Deep expertise in Kubernetes, including taints, cordons, node draining, and custom operators
  • Demonstrated experience with Go or Python in a production-level environment
  • Strong financial literacy and ability to model complex trade-offs between capacity reliability and cost
  • High tenacity and collaborative mindset
  • Experience with additional non-NVIDIA accelerators, such as AWS Trainium/Inferentia or AMD Instinct, is nice to have
  • Familiarity with multi-accelerator scheduling and cost/performance tradeoff modeling is nice to have
  • Prior experience partnering with model performance or ML systems teams is nice to have
Core Competencies

Demonstrates expertise in managing TPU pod fleets, including acquisition, allocation, and maintenance, while ensuring compliance with deployment scheduling rules. Proficient in designing capacity management systems and optimizing multi-cloud environments for cost and performance.

Highest-signal resume keywords
  • Google Cloud TPU Management
  • Kubernetes Expertise
  • Go or Python Programming
  • Capacity Management Systems
  • Financial Modeling
Hard Skills
  • TPU Pod Slicing
  • ICI Topology
  • JAX/XLA
  • Custom Operators
  • Node Draining
Soft Skills
  • Collaborative Mindset
  • High Tenacity
Industry Keywords
  • Hyperscaler
  • Accelerator Provider
  • Capacity Reliability
  • Cost/Performance Tradeoff
Tools & Technologies
  • Google Cloud
  • AWS
  • Azure
  • Multi-Cloud Systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Global Capacity Manager - TPU Focus
Global Capacity Manager - TPU Focus

Baseten • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Full medical coverage
Flexible PTO
+4
Global TPU Capacity Lead
Global TPU Capacity Lead

Baseten • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Full medical coverage
Flexible PTO
+4
Capacity Strategy – Operations
Capacity Strategy – Operations

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 170,000
Member of Technical Staff – AI Cloud Infrastructure
Member of Technical Staff – AI Cloud Infrastructure

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Global TPU Capacity Architect: Scale & Orchestrate
Global TPU Capacity Architect: Scale & Orchestrate

Baseten • United States

Remote
USD 180,000 - 240,000
Technical Program Manager - Data Center / HPC Infrastructure and Operations
Technical Program Manager - Data Center / HPC Infrastructure and Operations

CIeNET Technologies • Seattle (WA)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+5
Product Manager, Cloud TPUs
Product Manager, Cloud TPUs

Socket.dev • Sunnyvale (CA)

On-site
USD 163,000 - 236,000
Senior Technical Program Manager II, NPI, AI/ML (TPU) Systems
Senior Technical Program Manager II, NPI, AI/ML (TPU) Systems

Google • Sunnyvale (CA)

On-site
USD 240,000 - 333,000
Group Product Manager, High Performance Computing, Inbound, AI and Computing Infrastructure
Group Product Manager, High Performance Computing, Inbound, AI and Computing Infrastructure

Socket.dev • Sunnyvale (CA)

On-site
USD 240,000 - 333,000
Capacity Operations Manager
Capacity Operations Manager

Jobtailor • California (MO)

On-site
USD 180,000 - 230,000
Onsite SF/NY
Equity