Senior Software Engineer, Fleet-level ML Performance

Socket.dev

Sunnyvale (CA)

Hybrid

USD 174,000 - 252,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google Cloud's TPU Chip Architecture and Performance team analyzes ML workloads to balance performance and cost on next-generation TPU systems. You will work with model researchers, compiler developers, and systems engineers to drive architecture decisions and performance optimization.

The role focuses on end-to-end performance analysis, hardware/software co-design, and scalable estimation methods for AI infrastructure, with emphasis on power, reliability, and scheduling in a hyperscale

Qualifications

  • Bachelor's degree required in CS/EE/CE or related field.
  • 5 years of experience in systems architecture, power/performance trade-offs, data center or cloud hardware optimization.

Responsibilities

  • Perform fleet-level performance analysis of key ML workloads on future TPU systems using simulation tools.
  • Collaborate with ML researchers, compiler developers, systems engineers, and TPU architects to optimize performance across design space.
  • Design and implement a unified ML Accelerator Platform Performance Estimation Methodology using C++ and Python for scalable projections.
  • Define hardware/software requirements for future AI infrastructure with cross-functional teams.

Skills

RAS features knowledge
Deep learning workloads knowledge

Education

Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or related field
Master's degree in Electrical Engineering, Computer Engineering or Computer Science
PhD in Electrical Engineering, Computer Engineering or Computer Science

Job description

MINIMUM QUALIFICATIONS:
  • Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • 5 years of experience in systems architecture or computer architecture, power and performance trade-off analysis, or data center, cloud, infrastructure hardware optimization .
  • Experience with Reliability, Availability, and Serviceability (RAS) features, paradigms, or architecture.
PREFERRED QUALIFICATIONS:
  • Master's degree or PhD in Electrical Engineering, Computer Engineering or Computer Science, with an emphasis on computer architecture.
  • Knowledge of deep learning workloads, including embedding architectures and their hardware execution characteristics.
ABOUT THE JOB:

Google Cloud's mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action. In this role, you aren't just building technology; you're shaping the frontier of enterprise and driving the evolution of advanced models.

The TPU Chip Architecture and Performance team bridges Google's machine learning workloads and custom silicon architectures. Through hardware/software co-design, we define and shape the future TPU platforms required to meet Google's ambitious AI goals. In this role, you will conduct end-to-end performance analysis of critical ML workloads including Gemini, YouTube, Ads, and key 3P models on next-generation TPU systems. Using advanced simulation and projection methodologies, you will evaluate hardware features and software optimizations to balance performance and cost trade-offs. You will drive high-impact architecture decisions that maximize TPU efficiency while accounting for system-wide constraints such as power, reliability (RAS), and scheduling.

The AI and Infrastructure team is redefining what's possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google [https://www.google.com/about/careers/applications/benefits/].

RESPONSIBILITIES:
  • Perform fleet-level performance analysis of key ML workloads (e.g., Gemini) on future TPU systems using advanced simulation tools to evaluate hardware/software trade-offs and guide next-generation chip architecture.
  • Partner with teams across the ML stack including model researchers, compiler developers, systems engineers, and TPU architects to analyze and optimize performance across the design space.
  • Design and implement a unified ML Accelerator Platform Performance Estimation Methodology using C++ and Python to enable scalable performance and TCO projections across Google.
  • Collaborate cross-functionally with data center, hardware architecture, and framework teams to define key hardware and software requirements for future AI infrastructure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Co-Design Engineer, Google Cloud TPU
Performance Co-Design Engineer, Google Cloud TPU

Socket.dev • Sunnyvale (CA)

On-site
USD 192,000 - 278,000
Senior Software Engineer, Fleet-level ML Performance
Senior Software Engineer, Fleet-level ML Performance

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity grants
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Staff Software Engineer, TPU Performance
Senior Staff Software Engineer, TPU Performance

Socket.dev • Sunnyvale (CA)

On-site
USD 262,000 - 365,000
Equity
Benefits
Senior Performance Co-design Engineer, LLM Training
Senior Performance Co-design Engineer, LLM Training

Socket.dev • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Staff Software Engineer, TPU Performance
Staff Software Engineer, TPU Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Staff Software Engineer, ML Systems Co-Design
Staff Software Engineer, ML Systems Co-Design

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Staff Performance Codesign Engineer, TPU
Senior Staff Performance Codesign Engineer, TPU

Socket.dev • Sunnyvale (CA)

On-site
USD 240,000 - 333,000
Senior Performance Co-Design Engineer, TPU
Senior Performance Co-Design Engineer, TPU

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Senior Performance Co-Design Engineer, LLM Serving
Senior Performance Co-Design Engineer, LLM Serving

Socket.dev • Sunnyvale (CA)

On-site
USD 174,000 - 252,000