Scale & Infrastructure Engineer for Large-Scale AI

AI Breaking Wire

Mountain View, Northern (CA, KY)

Hybrid

USD 165,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Stock options
Healthcare coverage
PTO
Retirement match

Job summary

Google DeepMind is seeking a Research Engineer for the Infrastructure and Scale team. You will build foundational systems that train state‑of‑the‑art AI models and scale computation across large TPU clusters.

Responsibilities include designing high‑throughput distributed training frameworks, profiling performance bottlenecks on GPUs/TPUs, and collaborating with researchers to co‑design new features and system capabilities.

Qualifications

  • Strong systems background with distributed computing experience.
  • Proficiency in C++ and Python is required.
  • Experience with MPI, JAX or TensorFlow is preferred.

Responsibilities

  • Architect and optimize high-throughput distributed training frameworks for large models.
  • Profile and debug performance bottlenecks across hardware accelerators and software stacks.
  • Collaborate with researchers to co-design new algorithmic features and system capabilities.

Skills

Strong systems background
C++
Python
Distributed systems

Education

Bachelor's degree or higher in CS/EE or related

Tools

MPI
JAX
TensorFlow

Job description

Google DeepMind is seeking a Research Engineer for the Infrastructure and Scale team. You will build foundational systems that train state‑of‑the‑art AI models and scale computation across large TPU clusters.

Responsibilities include designing high‑throughput distributed training frameworks, profiling performance bottlenecks on GPUs/TPUs, and collaborating with researchers to co‑design new features and system capabilities.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Infrastructure and Scale
Research Engineer, Infrastructure and Scale

AI Breaking Wire • Mountain View (CA), Northern (KY)

Hybrid
USD 165,000 - 230,000
Stock options
Healthcare coverage
PTO
+1
Staff Software Engineer - Lead Large-Scale AI Infra
Staff Software Engineer - Lead Large-Scale AI Infra

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Health insurance
Dental insurance
Vision insurance
+8
Staff Software Engineer - AI & Large-Scale Infra
Staff Software Engineer - AI & Large-Scale Infra

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Optimization Research Engineer - Large-Scale ML/AI (Equity)
Optimization Research Engineer - Large-Scale ML/AI (Equity)

Google LLC • San Francisco (CA)

On-site
USD 147,000 - 210,000
Senior AI/ML Infrastructure Engineer for TPU Health
Senior AI/ML Infrastructure Engineer for TPU Health

Google • United States

On-site
USD 174,000 - 252,000
Health insurance
Dental insurance
Vision insurance
+8
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Senior AI/ML Infrastructure Engineer
Senior AI/ML Infrastructure Engineer

Google • New York (NY)

Hybrid
USD 194,000 - 253,000
Equity
Bonus target
AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Lead Infrastructure & AI Systems Engineer
Lead Infrastructure & AI Systems Engineer

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior Infrastructure Software Engineer (Large-Scale Systems)
Senior Infrastructure Software Engineer (Large-Scale Systems)

Google • Kirkland (WA)

On-site
USD 147,000 - 210,000
Health insurance
Dental insurance
Vision insurance
+8