GPU HPC Operations Manager - Reliability & Agile Leader

Northern Data Group

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Northern Data Group is seeking an Operations Engineering Manager to own the reliability of our GPU‑accelerated HPC infrastructure and lead a team of Operations Engineers. You will shape operational excellence, manage a growing team, and collaborate with Platform, Network, and Infrastructure teams while maturing scaled Agile practices.

You will drive proactive improvements, define operational metrics, and ensure efficient incident response, automation, and documentation across the Operations

Qualifications

  • 5+ years in infrastructure/operations, with 2+ years managing a technical team.
  • Advanced Linux administration in production, ideally at scale.
  • Experience with incident/problem management and third‑party support.
  • Hands‑on with automation and monitoring/observability tools.
  • Experience with Agile ways of working and scaled Agile frameworks.
  • Excellent communication and stakeholder management skills.

Responsibilities

  • Lead, coach, and develop an internal team of Operations Engineers.
  • Own the reliability, performance, and availability of GPU‑accelerated HPC infrastructure.
  • Oversee monitoring, incident trend analysis, and root cause analysis.
  • Define, track, and report on key operational metrics.
  • Drive automation initiatives and maintain runbooks and change control.

Skills

Leadership
Linux
Incident management
Automation
Monitoring
Python
Bash
CI/CD
DevOps tooling

Tools

Ansible
Grafana
Prometheus
NVIDIA GPUs
InfiniBand RDMA

Job description

Northern Data Group is seeking an Operations Engineering Manager to own the reliability of our GPU‑accelerated HPC infrastructure and lead a team of Operations Engineers. You will shape operational excellence, manage a growing team, and collaborate with Platform, Network, and Infrastructure teams while maturing scaled Agile practices.

You will drive proactive improvements, define operational metrics, and ensure efficient incident response, automation, and documentation across the Operations

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Operations Engineering Manager (m/f/d)
Operations Engineering Manager (m/f/d)

Northern Data Group • Greater London

Hybrid
GBP 90,000 - 130,000
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Data Center IT Ops Lead — GPU Clusters & Reliability
Data Center IT Ops Lead — GPU Clusters & Reliability

Nebius • Newport

On-site
GBP 55,000 - 85,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Triwill Group • United Kingdom

Remote
GBP 90,000 - 130,000
Data Center IT Ops Lead: GPU Clusters & High Availability
Data Center IT Ops Lead: GPU Clusters & High Availability

Nebius B.V. • Newport

On-site
GBP 60,000 - 75,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant • England

On-site
GBP 70,000 - 90,000
Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence
Senior Data Center Engineer — GPU & Linux Networking
Senior Data Center Engineer — GPU & Linux Networking

Pursuu • Manchester

Hybrid
GBP 40,000 - 70,000
Company events
Company pension
Free parking
+2
Remote Senior HPC Cluster Architect — GPU AI Infra Lead
Remote Senior HPC Cluster Architect — GPU AI Infra Lead

NexGen Cloud • Greater London

On-site
GBP 80,000 - 100,000
Competitive salary and annual discretionary bonus
Employee wellbeing benefits
25 days of holiday plus public holidays
+3
Senior Data & MLOps Engineer - GPU Reliability Platform
Senior Data & MLOps Engineer - GPU Reliability Platform

Coreweaveu • Greater London

On-site
GBP 120,000 - 180,000
Medical Insurance
Dental Insurance
Pension Contribution
+5
Hybrid Operations Engineer - Reliability & Incident Lead
Hybrid Operations Engineer - Reliability & Incident Lead

AGS • United Kingdom

Hybrid
GBP 51,000 - 69,000