AI Platform Engineer for Production ML & Kubernetes

United States Digital Space LLC

Greater London

On-site

GBP 75,000 - 95,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Comprehensive health coverage
Equity (RSUs)
Hybrid work model
Flexible time off
Paid parental leave

Job summary

Lightning AI, the company behind PyTorch Lightning, is hiring AIPlatform Support Engineers for the EMEA Customer Experience team. You will partner with ML engineers, diagnose failures in distributed systems, and guide customers through complex Kubernetes and GPU workloads in production.

We seek engineers who can diagnose performance bottlenecks, improve observability, and build runbooks. This hybrid London role requires in-office presence at least 2 days/week and supports large-scale AI

Qualifications

  • Strong software engineering and systems troubleshooting background.
  • Experience with Kubernetes and containerized environments.
  • Linux systems knowledge, including networking, storage, process management, and performance tuning.
  • Experience with cloud infrastructure and distributed systems.
  • Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.

Responsibilities

  • Partner directly with customer engineering teams running training and inference workloads in production.
  • Help customers diagnose and resolve complex distributed systems and ML infrastructure issues.
  • Act as a technical advisor during high impact incidents and platform degradation events.
  • Translate infrastructure level issues into actionable guidance for ML engineers.
  • Build credibility with customers through strong technical reasoning and clear communication.
  • Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems.
  • Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues.
  • Analyze logs, metrics, traces, and system behavior to isolate root causes.
  • Debug containerized workloads running across Kubernetes and bare metal GPU environments
  • Support customers scaling workloads across multi node GPU systems
  • Diagnose performance bottlenecks involving compute, memory, networking, or storage

Skills

Kubernetes
Linux
Cloud infrastructure
Distributed systems
Observability

Tools

Prometheus
Grafana
OpenTelemetry
PyTorch
CUDA
NCCL

Job description

Lightning AI, the company behind PyTorch Lightning, is hiring AIPlatform Support Engineers for the EMEA Customer Experience team. You will partner with ML engineers, diagnose failures in distributed systems, and guide customers through complex Kubernetes and GPU workloads in production.

We seek engineers who can diagnose performance bottlenecks, improve observability, and build runbooks. This hybrid London role requires in-office presence at least 2 days/week and supports large-scale AI

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform Engineer — Production & Reliability
AI Platform Engineer — Production & Reliability

Lightningai • Greater London

Hybrid
GBP 75,000 - 95,000
Health Coverage
Equity
Retirement Savings
+8
AI Platform Support Engineer (EMEA)
AI Platform Support Engineer (EMEA)

Lightningai • Greater London

Hybrid
GBP 75,000 - 95,000
Health Coverage
Equity
Retirement Savings
+8
AI Platform Support Engineer (EMEA)
AI Platform Support Engineer (EMEA)

United States Digital Space LLC • Greater London

Hybrid
GBP 75,000 - 95,000
Comprehensive health coverage
Equity (RSUs)
Hybrid work model
+2
Forward-Deployed Platform Engineer for AI Production
Forward-Deployed Platform Engineer for AI Production

Lightningai • Greater London

Hybrid
GBP 134,000 - 186,000
Health coverage
Equity and stock options
Retirement savings (401k/UK pension)
+8
Senior Platform Engineer, Core Backend
Senior Platform Engineer, Core Backend

Lightning AI • Greater London

Hybrid
GBP 134,000 - 186,000
Health coverage
Equity
Retirement benefits
+3
Platform Engineer (Forward Deployment)
Platform Engineer (Forward Deployment)

Lightningai • Greater London

Hybrid
GBP 134,000 - 186,000
Health coverage
Equity and stock options
Retirement savings (401k/UK pension)
+8
Backend Engineer - Go for Scalable AI Platform
Backend Engineer - Go for Scalable AI Platform

Lightningai • Greater London

Hybrid
GBP 134,000 - 186,000
Comprehensive health coverage
Equity grants
Retirement savings (US 401(k) / UK‑pN)
+3
Remote-Eligible Research Engineer, AI Systems
Remote-Eligible Research Engineer, AI Systems

Lightningai • Greater London

Hybrid
GBP 123,000 - 231,000
Health Coverage
Equity
Retirement Savings
+8
AI Infrastructure DevOps Engineer (Kubernetes + GPU)
AI Infrastructure DevOps Engineer (Kubernetes + GPU)

Carbon3.ai • Greater London

On-site
GBP 70,000 - 110,000
Senior Software Engineer, Core Platform
Senior Software Engineer, Core Platform

Socket.dev • Greater London

On-site
GBP 134,000 - 186,000
Health coverage
Equity
Retirement benefits
+3