Senior Infrastructure Operations Engineer - AI GPU Fleet

Academic Key

Greater London

Hybrid

GBP 110,000 - 155,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Lightning AI seeks a Senior Infrastructure Operations Engineer to operate and scale the world’s GPU fleet in its London office. You will own provisioning, monitoring, and incident response, driving reliability improvements across Linux, compute, and network layers.

You will build automation, partner with multiple engineering teams, and participate in on-call rotations to ensure production stability. Strong scripting, Kubernetes/Slurm experience, and a proactive mindset are required.

Qualifications

  • Strong experience operating and troubleshooting Linux-based production infrastructure at scale.
  • Experience troubleshooting complex infrastructure issues across compute, networking, storage, and OS.
  • Strong automation and scripting skills using Python, Go, Bash, Ansible, or similar tools.

Responsibilities

  • Operate and troubleshoot large-scale GPU and bare-metal infrastructure across Linux, compute, networking, storage, and cluster orchestration.
  • Serve as a technical escalation point for complex infrastructure issues, driving problems from initial investigation through resolution and root cause analysis.
  • Own operational workflows across the infrastructure lifecycle, including provisioning, configuration, validation, maintenance, remediation, and decommissioning.
  • Build automation and internal tooling that eliminates repetitive operational work and enables the infrastructure fleet to scale efficiently.
  • Identify recurring failure modes and partner with Infrastructure Engineering to build more reliable, repeatable, and automated systems.
  • Improve provisioning, monitoring, diagnostics, and operational processes across a rapidly growing infrastructure footprint.
  • Partner closely with Infrastructure Engineering, Network Engineering, Data Center Operations, Customer Experience, and Platform Engineering to resolve issues that cross team or system boundaries.
  • Participate in a distributed primary/secondary on-call rotation supporting production infrastructure.

Skills

Linux production infrastructure
Troubleshooting complex infra
Python
Go
Bash
Ansible
Kubernetes
Slurm
Monitoring & Observability
Networking fundamentals
Automation scripting

Tools

Prometheus/Grafana

Job description

Lightning AI seeks a Senior Infrastructure Operations Engineer to operate and scale the world’s GPU fleet in its London office. You will own provisioning, monitoring, and incident response, driving reliability improvements across Linux, compute, and network layers.

You will build automation, partner with multiple engineering teams, and participate in on-call rotations to ensure production stability. Strong scripting, Kubernetes/Slurm experience, and a proactive mindset are required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Infra Ops Engineer - Scale GPU Fleet (Hybrid)
Senior Infra Ops Engineer - Scale GPU Fleet (Hybrid)

United States Digital Space LLC • Greater London

Hybrid
GBP 116,000 - 128,000
Health coverage
Equity RSUs
401(k) matching
+3
Senior GPU Infra Operations Engineer – EMEA
Senior GPU Infra Operations Engineer – EMEA

Lightning • Greater London

Hybrid
GBP 90,000 - 150,000
Health coverage
Equity
401(k) matching
+3
Senior Infrastructure Engineer — GPU & Bare-Metal Automation
Senior Infrastructure Engineer — GPU & Bare-Metal Automation

Lightningai • Greater London

Hybrid
GBP 133,000 - 163,000
Health coverage
Equity
401k matching
+8
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

Academic Key • Greater London

Hybrid
GBP 110,000 - 155,000
Senior Infra Operations Engineer - Remote/UK Hybrid
Senior Infra Operations Engineer - Remote/UK Hybrid

lightningai • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Infrastructure Software Engineer - Remote & GPU Infra
Senior Infrastructure Software Engineer - Remote & GPU Infra

Lightning • Greater London

Hybrid
GBP 134,000 - 164,000
Health coverage
Equity: meaningful RSUs
401(k) retirement matching
+4
Senior Infra Software Engineer — Remote GPU/HPC
Senior Infra Software Engineer — Remote GPU/HPC

Precision Labs • Greater London

Hybrid
GBP 134,000 - 164,000
Comprehensive Health Coverage
Meaningful Equity
401(k) Matching (US)
+7
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

lightningai • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom
Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom

Lightning • Greater London

Hybrid
GBP 90,000 - 150,000
Health coverage
Equity
401(k) matching
+3
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

United States Digital Space LLC • Greater London

Hybrid
GBP 116,000 - 128,000
Health coverage
Equity RSUs
401(k) matching
+3