SRE Systems Engineering Manager, ML Compute

Google Inc.

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

37 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Google is hiring a Systems Engineering Manager for Site Reliability Engineering focusing on ML Compute in London. You will lead a team responsible for uptime, availability and performance of core services and drive automation across large-scale infrastructure.

Own end-to-end reliability, mentor engineers, manage on-call rotations across continents, and collaborate with product and platform teams to deliver scalable, fault-tolerant ML compute environments.

Qualifications

  • Bachelor's degree in Computer Science or a related technical field or equivalent practical experience.
  • 5 years of experience with programming in one or more programming languages.
  • 3 years of people management experience.
  • 3 years of experience leading projects and working with administration (e.g., filesystems, inodes, system calls) or networking (e.g., TCP/IP, routing, network topologies and hardware, SDN).

Responsibilities

  • Lead a team of software/systems engineers on projects for users and be directly responsible for uptime.
  • Own end-to-end availability and performance of key services and build automation to prevent problem recurrence. Automate response to all non-exceptional service conditions.
  • Lead by example, mentor the team and establish credibility through quality technical execution.
  • Manage on-call rotations across continents, using a follow-the-sun model.
  • Design, write and deliver software to improve the availability, scalability, latency and efficiency of Google's services.

Skills

Programming experience
People management
Project leadership
Networking basics
Decision making
Problem solving

Education

Bachelor's degree in Computer Science or related field

Job description

Google is hiring a Systems Engineering Manager for Site Reliability Engineering focusing on ML Compute in London. You will lead a team responsible for uptime, availability and performance of core services and drive automation across large-scale infrastructure.

Own end-to-end reliability, mentor engineers, manage on-call rotations across continents, and collaborate with product and platform teams to deliver scalable, fault-tolerant ML compute environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Compute SRE Lead: Scale, Uptime & Automation
ML Compute SRE Lead: Scale, Uptime & Automation

Google Inc. • City of Westminster

On-site
GBP 100,000 - 140,000
Systems Engineering Manager, Site Reliability Engineering, ML Compute
Systems Engineering Manager, Site Reliability Engineering, ML Compute

Google Inc. • Greater London

Hybrid
GBP 120,000 - 180,000
Systems Engineering Manager, Site Reliability Engineering, ML Compute
Systems Engineering Manager, Site Reliability Engineering, ML Compute

Google Inc. • City of Westminster

On-site
GBP 100,000 - 140,000
SRE Engineering Manager: Reliability & Automation
SRE Engineering Manager: Reliability & Automation

WeAreTechWomen • United Kingdom

On-site
GBP 110,000 - 150,000
Systems Engineering Manager, Site Reliability Engineering, ML Compute
Systems Engineering Manager, Site Reliability Engineering, ML Compute

Google • Greater London

On-site
GBP 120,000 - 170,000
Senior SRE Engineer — Cloud Reliability & Automation
Senior SRE Engineer — Cloud Reliability & Automation

Google Inc. • City of Westminster

On-site
GBP 90,000 - 140,000
Senior ML Inference & Model Serving Engineer
Senior ML Inference & Model Serving Engineer

Google • Greater London

On-site
GBP 160,000 - 200,000
Learning opportunities
Career growth
Senior SRE Engineer: Build Scalable, Fault-Tolerant Systems
Senior SRE Engineer: Build Scalable, Fault-Tolerant Systems

WeAreTechWomen • United Kingdom

On-site
GBP 90,000 - 130,000
SRE Manager: Reliability & Incident Leadership (Hybrid London)
SRE Manager: Reliability & Incident Leadership (Hybrid London)

Gravitas Recruitment Group (Global) Ltd • Greater London

Hybrid
GBP 75,000 - 100,000
Senior Data Scientist, Reliability Analytics & AI Tools
Senior Data Scientist, Reliability Analytics & AI Tools

Google Inc. • City of Westminster

On-site
GBP 90,000 - 130,000