Senior GPU Infra Operations Engineer – EMEA

Lightning

Greater London

Hybrid

GBP 90,000 - 150,000

Full time

22 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health coverage
Equity
401(k) matching
Unlimited PTO
Hybrid work
Office meals

Job summary

Lightning AI in London, UK, seeks a Senior Infrastructure Operations Engineer to operate and scale our GPU-based infrastructure across Linux, compute, networking, and storage. You will drive incident response, automate workflows, and partner with multiple engineering teams to improve reliability and scalability.

The role emphasizes ownership, collaboration, and building scalable, observable systems in a hybrid work environment.

Qualifications

  • Strong experience operating and troubleshooting Linux-based production infrastructure at scale.
  • Experience troubleshooting complex infrastructure issues across compute, networking, storage, and operating systems.
  • Strong automation and scripting skills using Python, Go, Bash, Ansible, or similar tools.
  • Experience with Kubernetes, Slurm, or other cluster and workload orchestration systems.
  • Experience using monitoring, observability, and telemetry systems to diagnose and troubleshoot production infrastructure.
  • Strong systems and networking fundamentals, with a track record of owning production issues through resolution.
  • Comfortable working in ambiguous environments and collaborating across engineering, operations, and customer-facing teams.

Responsibilities

  • Operate and troubleshoot large-scale GPU and bare-metal infrastructure across Linux, compute, networking, storage, and cluster orchestration.
  • Serve as a technical escalation point for complex infrastructure issues, driving problems from initial investigation through resolution and root cause analysis.
  • Own operational workflows across the infrastructure lifecycle, including provisioning, configuration, validation, maintenance, remediation, and decommissioning.
  • Build automation and internal tooling that eliminates repetitive operational work and enables the infrastructure fleet to scale efficiently.
  • Identify recurring failure modes and partner with Infrastructure Engineering to build more reliable, repeatable, and automated systems.
  • Improve provisioning, monitoring, diagnostics, and operational processes across a rapidly growing infrastructure footprint.
  • Partner closely with Infrastructure Engineering, Network Engineering, Data Center Operations, Customer Experience, and Platform Engineering to resolve issues that cross team or system boundaries.
  • Participate in a distributed primary/secondary on-call rotation supporting production infrastructure.

Skills

Linux at scale
Python scripting
Kubernetes
Networking basics
Observability

Tools

Ansible
Slurm

Job description

Lightning AI in London, UK, seeks a Senior Infrastructure Operations Engineer to operate and scale our GPU-based infrastructure across Linux, compute, networking, and storage. You will drive incident response, automate workflows, and partner with multiple engineering teams to improve reliability and scalability.

The role emphasizes ownership, collaboration, and building scalable, observable systems in a hybrid work environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer — GPU & Bare-Metal Automation
Senior Infrastructure Engineer — GPU & Bare-Metal Automation

United States Digital Space LLC • Greater London

Hybrid
GBP 133,000 - 163,000
Health coverage
Equity
401k matching
+8
Senior Infra Operations Engineer - Remote/UK Hybrid
Senior Infra Operations Engineer - Remote/UK Hybrid

lightningai • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Infra Software Engineer — Remote GPU/HPC
Senior Infra Software Engineer — Remote GPU/HPC

Precision Labs • Greater London

Hybrid
GBP 134,000 - 164,000
Comprehensive Health Coverage
Meaningful Equity
401(k) Matching (US)
+7
Senior Infrastructure Software Engineer - Remote & GPU Infra
Senior Infrastructure Software Engineer - Remote & GPU Infra

Lightning • Greater London

Hybrid
GBP 134,000 - 164,000
Health coverage
Equity: meaningful RSUs
401(k) retirement matching
+4
GPU & Compute Infra Engineer (Remote-friendly)
GPU & Compute Infra Engineer (Remote-friendly)

Lightningai • Greater London

Hybrid
GBP 136,000 - 166,000
Health coverage
Equity
401(k) matching
+7
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

lightningai • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom
Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom

Lightning • Greater London

Hybrid
GBP 90,000 - 150,000
Health coverage
Equity
401(k) matching
+3
Senior Infrastructure Software Engineer — Remote/Hybrid
Senior Infrastructure Software Engineer — Remote/Hybrid

Lightningai • Greater London

Hybrid
GBP 136,000 - 166,000
Health coverage
Equity (RSUs)
401(k) matching
+3
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Senior Data Center Engineer — GPU & Linux Networking
Senior Data Center Engineer — GPU & Linux Networking

Pursuu • Manchester

Hybrid
GBP 40,000 - 70,000
Company events
Company pension
Free parking
+2