Senior Infra Ops Engineer - Scale GPU Fleet (Hybrid)

United States Digital Space LLC

Greater London

Hybrid

GBP 116,000 - 128,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health coverage
Equity RSUs
401(k) matching
Unlimited PTO
Hybrid work model
Office meals

Job summary

Lightning AI in London is seeking an experienced Senior Infrastructure Operations Engineer to help operate and scale the infrastructure behind one of the world's largest GPU fleets. The role sits at the center of production reliability across Linux, bare-metal servers, GPUs, networking, storage, provisioning, and cluster orchestration.

You will own provisioning, monitoring, diagnostics, and on-call support, build automation, and partner with Infrastructure Engineering and Platform teams to

Qualifications

  • Strong Linux production infrastructure experience at scale.
  • Experience troubleshooting complex infra across compute, networking, storage, and OS.
  • Strong automation and scripting with Python, Go, Bash, Ansible, or similar.
  • Experience with Kubernetes, Slurm, or other cluster/workload schedulers.
  • Experience with monitoring, observability, and telemetry to diagnose issues.
  • Strong systems and networking fundamentals with ownership of production issues.
  • Comfortable collaborating across engineering, operations, and customer-facing teams.

Responsibilities

  • Operate and troubleshoot large-scale GPU and bare-metal infrastructure across Linux, compute, networking, storage, and cluster orchestration.
  • Serve as a technical escalation point for complex infrastructure issues, driving resolution and root cause analysis.
  • Own provisioning, configuration, validation, maintenance, remediation, and decommissioning workflows.
  • Build automation and internal tooling to scale the infrastructure fleet.
  • Identify recurring failures and work with Infra Eng to improve reliability and automation.
  • Improve provisioning, monitoring, diagnostics, and operational processes for a growing footprint.
  • Collaborate with Infrastructure Engineering, Network Engineering, Data Center Operations, Customer Experience, and Platform Engineering to resolve cross-team issues.
  • Participate in a distributed primary/secondary on-call rotation.

Skills

Linux ops
Automation scripts
Python scripting
Go scripting
Ansible
Kubernetes
Slurm
Monitoring
Telemetry
Networking basics
On-call readiness

Tools

PXE
BMC
IPMI
Redfish
iDRAC
NVIDIA GPUs
DCGM
InfiniBand
RoCE/RDMA
GPFS
Ceph

Job description

Lightning AI in London is seeking an experienced Senior Infrastructure Operations Engineer to help operate and scale the infrastructure behind one of the world's largest GPU fleets. The role sits at the center of production reliability across Linux, bare-metal servers, GPUs, networking, storage, provisioning, and cluster orchestration.

You will own provisioning, monitoring, diagnostics, and on-call support, build automation, and partner with Infrastructure Engineering and Platform teams to

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Infra Operations Engineer – EMEA
Senior GPU Infra Operations Engineer – EMEA

Lightning • Greater London

Hybrid
GBP 90,000 - 150,000
Health coverage
Equity
401(k) matching
+3
Senior Infra Operations Engineer - Remote/UK Hybrid
Senior Infra Operations Engineer - Remote/UK Hybrid

lightningai • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Infrastructure Engineer — GPU & Bare-Metal Automation
Senior Infrastructure Engineer — GPU & Bare-Metal Automation

Lightningai • Greater London

Hybrid
GBP 133,000 - 163,000
Health coverage
Equity
401k matching
+8
Senior Infra Software Engineer — Remote GPU/HPC
Senior Infra Software Engineer — Remote GPU/HPC

Precision Labs • Greater London

Hybrid
GBP 134,000 - 164,000
Comprehensive Health Coverage
Meaningful Equity
401(k) Matching (US)
+7
Senior Infrastructure Software Engineer - Remote & GPU Infra
Senior Infrastructure Software Engineer - Remote & GPU Infra

Lightning • Greater London

Hybrid
GBP 134,000 - 164,000
Health coverage
Equity: meaningful RSUs
401(k) retirement matching
+4
GPU & Compute Infra Engineer (Remote-friendly)
GPU & Compute Infra Engineer (Remote-friendly)

Lightningai • Greater London

Hybrid
GBP 136,000 - 166,000
Health coverage
Equity
401(k) matching
+7
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

United States Digital Space LLC • Greater London

Hybrid
GBP 116,000 - 128,000
Health coverage
Equity RSUs
401(k) matching
+3
Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom
Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom

Lightning • Greater London

Hybrid
GBP 90,000 - 150,000
Health coverage
Equity
401(k) matching
+3
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

lightningai • Greater London

Hybrid
GBP 90,000 - 130,000
GPU Infrastructure Lead: Scale, Certification, and Automation
GPU Infrastructure Lead: Scale, Certification, and Automation

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 140,000 - 170,000
Full Benefits