Member of Technical Staff (AI Inference Engineer)

Perplexity

Greater London

On-site

GBP 80,000 - 120,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

Perplexity is hiring an AI Inference Engineer to scale our inference engine behind every query. You will work on loading weights, scheduling requests, and managing KV-cache, with a stack of Rust, Python, CUDA, and CuTe DSL.

You will port CUDA kernels to CuTe DSL, optimize performance under tight latency and cost budgets, and build a robust Rust-based serving runtime to support growing traffic and model architectures.

Qualifications

  • 3+ years of professional software engineering experience in ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures and memory hierarchy.
  • Experience building production distributed systems under real load.
  • Comfortable working across languages: Rust for serving, Python for model code.

Responsibilities

  • Support new models including transformer-based retrieval, text generation, and multimodal models in inference infrastructure.
  • Migrate GPU kernels to CuTe DSL to run on GB200 today and Vera Rubin racks tomorrow.
  • Develop a Rust-based internal inference server to handle increasing traffic.
  • Profile and fix bottlenecks from network ingress through batching and GPU interleaving.
  • Build dashboards, alerts, and automated remediation to catch production regressions.

Skills

GPU programming
Performance optimisation
Distributed systems
Rust
Python
CUDA

Tools

CuTe DSL
CUDA kernels
Rust
Python
Triton

Job description

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.

Responsibilities
  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving.
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.
Who We're Looking For
  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.
  • You've built and operated production distributed systems under real load - ideally performance-critical ones.
  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
  • You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you.
Nice-to-have
  • ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
  • Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
  • Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
  • Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
  • Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads.
Qualifications
  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Final offer amounts are determined by multiple factors including experience and expertise.

Equity: In addition to the base salary, equity may be part of the total compensation package.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
AI Inference Engineer (Staff) - GPU ML Systems + Equity
AI Inference Engineer (Staff) - GPU ML Systems + Equity

Perplexity • Greater London

On-site
GBP 80,000 - 120,000
Equity
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Member of Technical Staff - Mid-Training Infra
Member of Technical Staff - Mid-Training Infra

Reflection • Greater London

On-site
GBP 70,000 - 100,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Inference System & Performance - Member of Technical Staff
Inference System & Performance - Member of Technical Staff

Callosum • Greater London

On-site
GBP 101,000 - 192,000
Competitive salary
Equity & Ownership
Private healthcare
+2
Senior Performance Engineer | AI Infrastructure | Cambridge
Senior Performance Engineer | AI Infrastructure | Cambridge

Pure Resourcing Solutions • Dry Drayton

Hybrid
GBP 90,000 - 120,000
AI Inference Engineer
AI Inference Engineer

Fuse Energy, LLC • Greater London

Hybrid
GBP 120,000 - 180,000
Equity
Biannual bonus
Fully expensed tech
+2
Senior Performance Engineer | AI Infrastructure | Cambridge |
Senior Performance Engineer | AI Infrastructure | Cambridge |

Pure Resourcing Solutions • Cambridge

Hybrid
GBP 90,000 - 120,000
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)

Pure Resourcing Solutions Limited • Dry Drayton

Hybrid
GBP 90,000 - 120,000
Competitive salary
Pension
Hybrid from Cambridge office
+1
Research Engineer (Inference & Serving)
Research Engineer (Inference & Serving)

Axiōma Search • Greater London

On-site
GBP 90,000 - 120,000