Research Scientist/Engineer, Frontier Reasoning, DeepMind

DeepMind Technologies Limited

Mountain View (CA)

On-site

USD 207,000 - 300,000

Full time

21 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Bonus target
Benefits package

Job summary

DeepMind Technologies Limited invites applications for a senior researcher/engineer on the PRISM team. You will work across the full research and engineering lifecycle, developing distributed post-training infrastructure and algorithms that enable Gemini models to solve complex multistep problems autonomously.

We pursue breakthroughs and production-ready capabilities for Gemini and Gemma, with a focus on scalable reasoning, RL scaling, and responsible AI.

Qualifications

  • Bachelor's or Master's in a quantitative field with practical experience.
  • 4+ years building and scaling ML models with deep learning frameworks.
  • Experience in RL, SFT/RLHF/RLAIF, or related agentic areas.

Responsibilities

  • Operate across the full research-and-engineering lifecycle of frontier reasoning and agentic systems.
  • Address unsolved problems in agentic reasoning, turning early prototypes into hardened production features for Gemini releases.
  • Architect and optimize distributed post-training pipelines and agent-environment simulation loops across thousands of accelerators.
  • Design rigorous experiments and failure analyses to isolate performance bottlenecks and communicate findings.
  • Drive technical excellence by maintaining high code quality and architectural health across shared reinforcement learning and modeling codebases.

Skills

ML frameworks
Reinforcement Learning
Distributed systems
Python

Education

Bachelor's/Master's in CS/Math/Physics

Tools

JAX
PyTorch
TensorFlow

Job description

Note: By applying to this position you will have an opportunity to share your preferred working location from the following: London, UK; Mountain View, CA, USA; New York, NY, USA.

Minimum qualifications:
  • Bachelor's or Master's degree in Computer Science, Mathematics, Physics, a related quantitative field, or equivalent practical experience.
  • 4 years of experience building, scaling, and debugging machine learning models using deep learning frameworks (e.g., JAX, PyTorch, or TensorFlow).
  • Experience in at least one core area: Reinforcement Learning (RL), Post-Training (SFT/RLHF/RLAIF), Agentic Tool-Use, or Inference-Time Search.
Preferred qualifications:
  • PhD in Computer Science, Machine Learning, Physics, or a related quantitative field.
  • Experience designing asynchronous agent-environment simulation loops or large distributed post-training pipelines.
  • Experience prototyping new hypotheses quickly while keeping shared codebases clean, robust, and production-grade.
About the job

At DeepMind, the Planning, Reasoning, Inference and Structured Models (PRISM) team brings together researchers and engineers to advance the frontiers of AI reasoning and autonomous agentic systems. We reject the false tradeoff between research and execution, pursuing breakthroughs on open AI challenges while embedding directly into core teams to land those capabilities in production. Our work powers Gemini and Gemma by developing core reasoning capabilities and RL scaling for Gemini 3, and leading Gemma 270M, including multiagent Gemini capabilities. We deliver critical contributions to AI Grand Challenges such as our gold medal winning IMO 2025 effort, drive product innovations like deep think mode and agentic inference scaling in antigravity, and lead Alphabet wide initiatives including AI for Science and Project Big Sleep. In this role, you will operate across the full research and engineering lifecycle, developing distributed post training infrastructure and algorithms that enable Gemini models to solve complex, multistep problems autonomously. Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority. We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort. Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits.

Responsibilities

Learn more about benefits at Google .

  • Operate across the full research-and-engineering lifecycle of frontier reasoning and agentic systems.
  • Address unsolved problems in agentic reasoning, turning early exploratory prototypes into hardened production features for Gemini releases.
  • Architect and optimize distributed post-training pipelines and agent-environment simulation loops across thousands of accelerators.
  • Design rigorous experiments and failure analyses to isolate performance bottlenecks and communicate findings through clear write-ups.
  • Drive technical excellence by maintaining high code quality and architectural health across shared reinforcement learning and modeling codebases.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Post-training Agentic Research Scientist, DeepMind
Post-training Agentic Research Scientist, DeepMind

Google DeepMind • Mountain View (CA)

On-site
USD 174,000 - 252,000
Bonus target 15%
Equity
Benefits
Post-training Agentic Research Scientist, DeepMind
Post-training Agentic Research Scientist, DeepMind

Google DeepMind • New York (NY)

On-site
USD 174,000 - 252,000
Post-training Agentic Research Scientist, DeepMind
Post-training Agentic Research Scientist, DeepMind

Google Inc. • Mountain View (CA), Northern (KY)

On-site
USD 190,000 - 236,000
Equity
Benefits
Senior Research Engineer, Agentic Data and Tooling, DeepMind
Senior Research Engineer, Agentic Data and Tooling, DeepMind

Google • Ionia (NY)

On-site
USD 174,000 - 252,000
Post-training Agentic Research Scientist, DeepMind
Post-training Agentic Research Scientist, DeepMind

Google • New York (NY)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Research Engineer, Advancing Agent Quality, DeepMind
Research Engineer, Advancing Agent Quality, DeepMind

Google DeepMind • Mountain View (CA)

On-site
USD 174,000 - 252,000
Gemini Post-training Software Engineer, DeepMind
Gemini Post-training Software Engineer, DeepMind

Google DeepMind • New York (NY)

On-site
USD 174,000 - 252,000
Equity
Gemini Post-training Software Engineer, DeepMind
Gemini Post-training Software Engineer, DeepMind

Google DeepMind • Mountain View (CA)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Gemini Post-training Software Engineer, DeepMind
Gemini Post-training Software Engineer, DeepMind

Google • New York (NY)

On-site
USD 174,000 - 252,000
Staff Research Scientist/Engineer, Agentic Post Training, DeepMind
Staff Research Scientist/Engineer, Agentic Post Training, DeepMind

DeepMind Technologies Limited • New York (NY)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid time off
+4