Platform Engineer: GPU Inference & Training

Perplexity

Greater London

On-site

GBP 100,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Perplexity is building a dedicated platform team to own and simplify GPU-driven training and inference workloads. You will design and run a self-serve compute platform, operate a GPU fleet across cloud providers, and implement scheduling, fault tolerance, and observability to keep workloads healthy and predictable.

You will own Kubernetes-based GPU orchestration, write operators and CRDs, and collaborate with inference and cloud infra engineers to align on architecture and roadmap.

Qualifications

  • Deep Kubernetes experience - custom operators, CRDs, and multi-cluster federation.
  • You've managed GPU clusters at scale: NVIDIA hardware, CUDA, and networking that makes them fast (InfiniBand or RoCE).
  • You've orchestrated compute across multiple clouds (CoreWeave, AWS, GCP, or similar) and understand how different each one really is.
  • Strong distributed systems fundamentals: scheduling, resource allocation, and fault tolerance under load.
  • You write infrastructure and systems-level code in Go, Rust or C++.
  • You've supported both long-running training jobs and high-availability inference services, and you know why they pull infrastructure in opposite directions.
  • You own problems end-to-end and do well when the path forward isn't laid out for you.

Responsibilities

  • Build a self-serve compute platform. Design and own systems that let teams launch training jobs and run inference services without managing GPU provisioning or provider-specific infra.
  • Operate the GPU fleet: provisioning, lifecycle management, reliability, and capacity integration across providers.
  • Solve for GPU scarcity: scheduling and placement logic to maximize utilization across providers under real constraints.
  • Support long-running training and production inference workloads on the same fleet with high availability.
  • Own Kubernetes for GPU orchestration: write operators and CRDs, manage clusters across providers.
  • Build fault tolerance, autoscaling, and observability to survive node loss or capacity shifts without manual intervention.
  • Set technical direction across teams to shape platform architecture and roadmap.

Skills

Kubernetes - operators & CRDs
GPU cluster management
Multi-cloud orchestration
Distributed systems fundamentals
Go
Rust
C++

Tools

CUDA
InfiniBand
RoCE
Prometheus
Grafana
Weights & Biases

Job description

Perplexity is building a dedicated platform team to own and simplify GPU-driven training and inference workloads. You will design and run a self-serve compute platform, operate a GPU fleet across cloud providers, and implement scheduling, fault tolerance, and observability to keep workloads healthy and predictable.

You will own Kubernetes-based GPU orchestration, write operators and CRDs, and collaborate with inference and cloud infra engineers to align on architecture and roadmap.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

Perplexity • Greater London

On-site
GBP 100,000 - 160,000
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Platform Engineer - AI Infra & Large-Scale GPU Systems
Platform Engineer - AI Infra & Large-Scale GPU Systems

Ineffable • Greater London

On-site
GBP 90,000 - 110,000
AI Inference Engineer | GPU-Scale Rust/Python | Equity
AI Inference Engineer | GPU-Scale Rust/Python | Equity

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
Equity options
Competitive compensation
Staff Inference Platform Engineer — Low-Latency GPU, Kubernetes
Staff Inference Platform Engineer — Low-Latency GPU, Kubernetes

CoreWeave Europe • Greater London

On-site
GBP 120,000 - 190,000
Medical Insurance
Pension Plan
Life Insurance
+1
Compute Platform Engineer — Multi-Cloud GPU & Kubernetes
Compute Platform Engineer — Multi-Cloud GPU & Kubernetes

Reflection AI • Greater London

On-site
GBP 80,000 - 120,000
Platform Engineering Lead: GPU Infra & Kubernetes
Platform Engineering Lead: GPU Infra & Kubernetes

Volta • Greater London

On-site
GBP 120,000 - 180,000
Equity in Volta
Retirement benefits
Health & wellbeing benefits
+1
Platform Engineer – Scale AI Infra & GPU Orchestration
Platform Engineer – Scale AI Infra & GPU Orchestration

Ineffable Intelligence • Greater London

On-site
GBP 70,000 - 110,000
HPC & AI Platform Engineer (GPU/Networking)
HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000