AI Platform Engineer — Production & Reliability

Lightningai

Greater London

Hybrid

GBP 75,000 - 95,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health Coverage
Equity
Retirement Savings
Unlimited PTO
Winter Break
Parental Leave
Learning Allowance
Wellness Stipend
Sabbatical
Flexible Schedule
In-Office Meals

Job summary

Lightning AI is hiring AI Platform Support Engineers to join our EMEA Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production. This role is hybrid out of our London office, with at least 2 in-office days per week and no visa sponsorship available at this time.

We operate two shifts (9AM–7PM CET/CEST) across Saturday–Tuesday and Thursday–Sunday, and you’ll partner closely with

Qualifications

  • Strong software engineering and systems troubleshooting background.
  • Experience with Kubernetes and containerized environments.
  • Linux systems knowledge including networking, storage, and performance tuning.
  • Experience with cloud infrastructure and distributed systems.
  • Observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.

Responsibilities

  • Partner directly with customer engineering teams running training and inference workloads in production.
  • Help customers diagnose and resolve complex distributed systems and ML infrastructure issues.
  • Act as a technical advisor during high impact incidents and platform degradation events.
  • Translate infrastructure level issues into actionable guidance for ML engineers.
  • Build credibility with customers through strong technical reasoning and clear communication.
  • Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems.
  • Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues.
  • Analyze logs, metrics, traces, and system behavior to isolate root causes.
  • Debug containerized workloads running across Kubernetes and bare metal GPU environments.
  • Support customers scaling workloads across multi node GPU systems.
  • Diagnose performance bottlenecks involving compute, memory, networking, or storage.
  • Identify recurring patterns across customer issues and drive long term reliability improvements.
  • Contribute to post incident reviews and operational improvements.
  • Build internal tooling, automation, documentation, and runbooks.
  • Partner closely with infrastructure, networking, and platform engineering teams.
  • Help improve observability, operational visibility, and troubleshooting workflows.
  • Improve the customer experience through better processes and technical guidance.

Skills

Kubernetes
Linux
Cloud infrastructure
Prometheus
Grafana
OpenTelemetry
Distributed systems
Python scripting

Tools

Ray
Kubeflow
Slurm
InfiniBand
RDMA
PyTorch
CUDA
NCCL

Job description

Lightning AI is hiring AI Platform Support Engineers to join our EMEA Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production. This role is hybrid out of our London office, with at least 2 in-office days per week and no visa sponsorship available at this time.

We operate two shifts (9AM–7PM CET/CEST) across Saturday–Tuesday and Thursday–Sunday, and you’ll partner closely with

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform Engineer for Production ML & Kubernetes
AI Platform Engineer for Production ML & Kubernetes

United States Digital Space LLC • Greater London

On-site
GBP 75,000 - 95,000
Comprehensive health coverage
Equity (RSUs)
Hybrid work model
+2
AI Platform Support Engineer (EMEA)
AI Platform Support Engineer (EMEA)

Lightningai • Greater London

Hybrid
GBP 75,000 - 95,000
Health Coverage
Equity
Retirement Savings
+8
AI Platform Support Engineer (EMEA)
AI Platform Support Engineer (EMEA)

United States Digital Space LLC • Greater London

Hybrid
GBP 75,000 - 95,000
Comprehensive health coverage
Equity (RSUs)
Hybrid work model
+2
Forward-Deployed Platform Engineer for AI Production
Forward-Deployed Platform Engineer for AI Production

Lightningai • Greater London

Hybrid
GBP 134,000 - 186,000
Health coverage
Equity and stock options
Retirement savings (401k/UK pension)
+8
Backend Engineer - Go for Scalable AI Platform
Backend Engineer - Go for Scalable AI Platform

Lightningai • Greater London

Hybrid
GBP 134,000 - 186,000
Comprehensive health coverage
Equity grants
Retirement savings (US 401(k) / UK‑pN)
+3
Platform Engineer (Forward Deployment)
Platform Engineer (Forward Deployment)

Lightningai • Greater London

Hybrid
GBP 134,000 - 186,000
Health coverage
Equity and stock options
Retirement savings (401k/UK pension)
+8
Remote-Eligible Research Engineer, AI Systems
Remote-Eligible Research Engineer, AI Systems

Lightningai • Greater London

Hybrid
GBP 123,000 - 231,000
Health Coverage
Equity
Retirement Savings
+8
AI Platform Support Engineer II (Hybrid London)
AI Platform Support Engineer II (Hybrid London)

Weights & Biases • Greater London

On-site
GBP 56,000 - 75,000
Medical insurance
Equity awards
Discretionary bonus
+1
Senior AI Infra Engineer - ML Platform & GPU Systems
Senior AI Infra Engineer - ML Platform & GPU Systems

Salient Group • Greater London

Hybrid
GBP 120,000 - 180,000
Equity grant
AI Platform Engineer: Automation, Resilience & Scale
AI Platform Engineer: Automation, Resilience & Scale

M&G • City of Edinburgh

On-site
GBP 65,000 - 95,000