Research Scientist/Engineer, Frontier Reasoning, DeepMind

AI Chopping Block, Inc.

Greater London

Hybrid

GBP 156,000 - 226,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Google DeepMind is seeking researchers and engineers to advance AI reasoning and autonomous agentic systems within the PRISM team. You will work across the full life cycle, building distributed post-training infrastructure to enable Gemini models to solve complex, multi-step problems automatically.

The position targets professionals with 4+ years in ML model development using DL frameworks and RL, SFT/RLHF/RLAIF experience, and a track record in designing scalable, production-grade systems.

Qualifications

  • Bachelor's or Master’s degree in CS, Math, Physics or related field, or equivalent practical experience.
  • 4 years of experience building, scaling, and debugging ML models using DL frameworks (JAX, PyTorch, TF).
  • Experience in RL, Post-Training (SFT/RLHF/RLAIF), Agentic Tool-Use, or Inference-Time Search.

Responsibilities

  • Operate across the full lifecycle, building distributed post-training infrastructure for Gemini multi-step problems.
  • Bridge research to production by turning prototypes into production features for Gemini releases.
  • Build and scale systems: design distributed post-training pipelines and simulations across accelerators.
  • Run ablations: design experiments and analyze failures; communicate findings clearly.
  • Maintain high code quality and architectural health across RL and modeling codebases.

Skills

Machine learning
Deep learning frameworks
Reinforcement Learning
Post-Training
Agentic Tool-Use
Inference-Time Search

Education

Bachelor's or Master's degree in CS/Math/Physics
PhD in CS/ML/Physics

Tools

JAX
PyTorch
TensorFlow

Job description

At Google DeepMind, the PRISM (Planning, Reasoning, Inference & Structured Models) team brings together researchers and engineers to advance the frontiers of AI reasoning and autonomous agentic systems. We reject the false tradeoff between research and execution, pursuing breakthroughs on open AI challenges while embedding directly into core teams to land those capabilities in production.

Our work powers Gemini & Gemma—developing core reasoning, multi-agent capabilities, RL and inference scaling in Gemini, and spearheading Gemma 270M. We deliver critical contributions to AI Grand Challenges (such as our gold medal-winning IMO 2025 effort), drive product innovations like 'Deep Think' mode and agentic inference scaling in antigravity, and contribute to Alphabet-wide initiatives including AI for Science and Project Big Sleep.

You will operate across the full research-and-engineering lifecycle, developing distributed post-training infrastructure and algorithms that enable Gemini models to solve complex, multi-step problems autonomously.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

  • You will operate across the full research-and-engineering lifecycle of frontier reasoning and agentic systems.
  • Bridge research to production:Tackleunsolved problems in agentic reasoning, turning early exploratory prototypes into hardened production features for Gemini releases.
  • Build and scale systems: Architect and optimize distributed post-training pipelines and agent-environment simulation loops across thousands of accelerators.
  • Run scientific ablations: Design rigorous experiments and failure analyses to isolate performance bottlenecks and communicate findings through clear write-ups.
  • Drive technical excellence: Maintain high code quality and architectural health across shared Reinforcement Learning and modeling codebases.
Minimum qualifications:
  • Bachelor's or Master's degree in Computer Science, Mathematics, Physics, a related quantitative field, or equivalent practical experience.
  • 4 years of experience building, scaling, and debugging machine learning models using deep learning frameworks (e.g., JAX, PyTorch, or TensorFlow).
  • Experience in at least one core area: Reinforcement Learning (RL), Post-Training (SFT/RLHF/RLAIF), Agentic Tool-Use, or Inference-Time Search.
Preferred qualifications:
  • PhD in Computer Science, Machine Learning, Physics, or a related quantitative field.
  • Experience designing asynchronous agent-environment simulation loops or large distributed post-training pipelines.
  • Experience prototyping new hypotheses quickly while keeping shared codebases clean, robust, and production-grade.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist: Agentic AI & Reasoning in Production
Research Scientist: Agentic AI & Reasoning in Production

AI Chopping Block, Inc. • Greater London

Hybrid
GBP 156,000 - 226,000
Research Scientist, Gemini Safety and Behavior, DeepMind
Research Scientist, Gemini Safety and Behavior, DeepMind

DeepMind Technologies Limited • Greater London

On-site
GBP 120,000 - 180,000
Equity
Benefits
Research Scientist, Robotics, DeepMind
Research Scientist, Robotics, DeepMind

Software Careers • Greater London

On-site
GBP 90,000 - 140,000
Research Scientist, Robotics RL, DeepMind
Research Scientist, Robotics RL, DeepMind

Google LLC • Greater London

Hybrid
GBP 120,000 - 180,000
Research Scientist, Robotics RL, DeepMind
Research Scientist, Robotics RL, DeepMind

DeepMind Technologies Limited • Greater London

On-site
GBP 90,000 - 130,000
Software Engineer, Model Inference, DeepMind
Software Engineer, Model Inference, DeepMind

DeepMind Technologies Limited • Greater London

Hybrid
GBP 140,000 - 200,000
Research Scientist, Robotics RL, DeepMind
Research Scientist, Robotics RL, DeepMind

Google Inc. • Greater London

On-site
GBP 120,000 - 150,000
Research Scientist, Robotics Pre-Training and Data Quality, DeepMind
Research Scientist, Robotics Pre-Training and Data Quality, DeepMind

Google LLC • Greater London

On-site
GBP 110,000 - 150,000
Research Scientist, Robotics RL, DeepMind
Research Scientist, Robotics RL, DeepMind

Google • Greater London

On-site
GBP 120,000 - 180,000
Research Scientist, Multisensor Robotics, DeepMind
Research Scientist, Multisensor Robotics, DeepMind

Google LLC • Greater London

Hybrid
GBP 90,000 - 150,000