AI Inference Engineer (Staff) - GPU ML Systems + Equity

Perplexity

Greater London

On-site

GBP 80,000 - 120,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

Perplexity is hiring an AI Inference Engineer to scale our inference engine behind every query. You will work on loading weights, scheduling requests, and managing KV-cache, with a stack of Rust, Python, CUDA, and CuTe DSL.

You will port CUDA kernels to CuTe DSL, optimize performance under tight latency and cost budgets, and build a robust Rust-based serving runtime to support growing traffic and model architectures.

Qualifications

  • 3+ years of professional software engineering experience in ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures and memory hierarchy.
  • Experience building production distributed systems under real load.
  • Comfortable working across languages: Rust for serving, Python for model code.

Responsibilities

  • Support new models including transformer-based retrieval, text generation, and multimodal models in inference infrastructure.
  • Migrate GPU kernels to CuTe DSL to run on GB200 today and Vera Rubin racks tomorrow.
  • Develop a Rust-based internal inference server to handle increasing traffic.
  • Profile and fix bottlenecks from network ingress through batching and GPU interleaving.
  • Build dashboards, alerts, and automated remediation to catch production regressions.

Skills

GPU programming
Performance optimisation
Distributed systems
Rust
Python
CUDA

Tools

CuTe DSL
CUDA kernels
Rust
Python
Triton

Job description

Perplexity is hiring an AI Inference Engineer to scale our inference engine behind every query. You will work on loading weights, scheduling requests, and managing KV-cache, with a stack of Rust, Python, CUDA, and CuTe DSL.

You will port CUDA kernels to CuTe DSL, optimize performance under tight latency and cost budgets, and build a robust Rust-based serving runtime to support growing traffic and model architectures.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 80,000 - 120,000
Equity
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
Senior AI Infrastructure Performance Engineer - GPU/Inference
Senior AI Infrastructure Performance Engineer - GPU/Inference

Pure Resourcing Solutions • Cambridge

Hybrid
GBP 90,000 - 120,000
Senior Performance Engineer — AI Inference & Systems Optimizer
Senior Performance Engineer — AI Inference & Systems Optimizer

CommonAI CIC • Cambridge

On-site
GBP 70,000 - 100,000
Competitive salary
Pension
Professional development
+2
Founding AI Inference Engineer – Scale & Serving Expert
Founding AI Inference Engineer – Scale & Serving Expert

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Junior Performance Engineer, AI Infra - Hybrid Cambridge
Junior Performance Engineer, AI Infra - Hybrid Cambridge

Pure Resourcing Solutions Limited • Linton

Hybrid
GBP 42,000 - 70,000
Pension
Hybrid from Cambridge office
Senior DL Inference Engineer — GPU-Accelerated AI at Scale
Senior DL Inference Engineer — GPU-Accelerated AI at Scale

NVIDIA • United Kingdom

On-site
GBP 110,000 - 160,000
Competitive salaries
Extensive benefits package
Diversity & inclusion
Staff Software Engineer, Kubernetes-native GPU Inference
Staff Software Engineer, Kubernetes-native GPU Inference

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Edge AI Inference Engineer: Kernel & Serving
Edge AI Inference Engineer: Kernel & Serving

Tether • United Kingdom

Remote
GBP 80,000 - 100,000