Research Engineer, Infrastructure and Scale

AI Breaking Wire

Mountain View, Northern (CA, KY)

Hybrid

USD 165,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
Healthcare coverage
PTO
Retirement match

Job summary

Google DeepMind is seeking a Research Engineer for the Infrastructure and Scale team. You will build foundational systems that train state‑of‑the‑art AI models and scale computation across large TPU clusters.

Responsibilities include designing high‑throughput distributed training frameworks, profiling performance bottlenecks on GPUs/TPUs, and collaborating with researchers to co‑design new features and system capabilities.

Qualifications

  • Strong systems background with distributed computing experience.
  • Proficiency in C++ and Python is required.
  • Experience with MPI, JAX or TensorFlow is preferred.

Responsibilities

  • Architect and optimize high-throughput distributed training frameworks for large models.
  • Profile and debug performance bottlenecks across hardware accelerators and software stacks.
  • Collaborate with researchers to co-design new algorithmic features and system capabilities.

Skills

Strong systems background
C++
Python
Distributed systems

Education

Bachelor's degree or higher in CS/EE or related

Tools

MPI
JAX
TensorFlow

Job description

About the Role

Google DeepMind is looking for a Research Engineer to join our core Infrastructure and Scale team. You will build the foundational systems that train some of the world's most advanced AI models. Your work will directly impact how we scale computation across massive TPU clusters, driving the next wave of scientific and algorithmic breakthroughs.

Responsibilities
  • Architect and optimize high-throughput distributed training frameworks for large language models and multimodal systems.
  • Profile and debug complex performance bottlenecks across hardware accelerators, interconnects, and software stacks.
  • Collaborate with researchers to co-design new algorithmic features and system-level capabilities.
Requirements
  • B.S., M.S., or Ph.D. in Computer Science, Electrical Engineering, or a related technical discipline.
  • Strong systems background with proficiency in C++ and Python.
  • Hands-on experience with distributed systems, MPI, JAX, or TensorFlow.
  • Experience scaling workloads across large clusters of GPUs or TPUs.
Benefits
  • Competitive compensation package including salary, bonus, and stock units.
  • Comprehensive healthcare and wellness programs.
  • Generous paid time off and retirement matching programs.
  • Access to cutting-edge computing resources and continuous learning opportunities.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Scale & Infrastructure Engineer for Large-Scale AI
Scale & Infrastructure Engineer for Large-Scale AI

AI Breaking Wire • Mountain View (CA), Northern (KY)

Hybrid
USD 165,000 - 230,000
Stock options
Healthcare coverage
PTO
+1
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Research Engineer, GenAI, Info Task, DeepMind
Research Engineer, GenAI, Info Task, DeepMind

Google DeepMind • Cambridge (MA)

On-site
USD 174,000 - 252,000
AI Research Engineer
AI Research Engineer

AI Breaking Wire • Northern (KY), New York (NY)

Hybrid
USD 210,000 - 310,000
Competitive salary
Health insurance
Dental and vision insurance
+2
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

Google DeepMind • Mountain View (CA)

Hybrid
USD 230,000 - 290,000
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Engineer, Responsible Frontier AI Research, DeepMind
Research Engineer, Responsible Frontier AI Research, DeepMind

Google DeepMind • New York (NY)

On-site
USD 207,000 - 300,000
Research Engineer, GenAI, Information Tasks, DeepMind
Research Engineer, GenAI, Information Tasks, DeepMind

Google Inc. • Cambridge (MA)

On-site
USD 174,000 - 252,000
Research Engineer, GenAI, Info Task, DeepMind
Research Engineer, GenAI, Info Task, DeepMind

Socket.dev • Cambridge (MA)

On-site
USD 174,000 - 252,000
Equity
Bonus plan (15% target)
Benefits
Software Engineer, Machine Learning Infrastructure
Software Engineer, Machine Learning Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

On-site
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1