NLP Performance Engineer: Scale LLM Inference for Research

Barlowe LLP

Greater London

On-site

GBP 90,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation + bonus
Lunch provided
35 days annual leave
9% pension contributions
Casual dress code
Healthcare & life assurance
Cycle-to-work
Monthly company events

Job summary

G-Research is seeking an exceptional NLP Performance Engineer to own large-scale LLM inference performance within our NLP Engineering team at the London HQ. You will profile, optimise and implement techniques to maximise inference throughput and cost-efficiency, collaborating with researchers and infrastructure engineers to shape scalable tooling and compute stacks.

The role blends high-impact hands-on work with systems design, offering a competitive package and a dynamic research-driven

Qualifications

  • Bachelor’s, Master’s or PhD in computer science, or equivalent experience.
  • Proven experience profiling, benchmarking and optimising large-scale LLM inference workloads.
  • A scientific, evidence-led approach to performance optimisation with rigorous benchmarking.

Responsibilities

  • Profiling, benchmarking and optimising large-scale LLM inference workloads across compute infrastructure.
  • Ensuring efficient deployment of latest models across GPU architectures.
  • Designing and implementing inference optimisations while maintaining output quality.
  • Developing reference implementations, libraries and tooling for NLP workloads.
  • Collaborating with researchers and engineers to design optimised solutions.
  • Working with systems and platform teams to evolve the compute stack.

Skills

LLM inference
Profiling
Benchmarking
Transformer inference
Python
CUDA
PyTorch
Model parallelism
Quantisation
Speculative decoding
SGLang
vLLM
TensorRT-LLM
Communication skills
Software engineering

Education

Bachelor/Master/PhD in CS

Tools

PyTorch ecosystem
LLM serving frameworks

Job description

G-Research is seeking an exceptional NLP Performance Engineer to own large-scale LLM inference performance within our NLP Engineering team at the London HQ. You will profile, optimise and implement techniques to maximise inference throughput and cost-efficiency, collaborating with researchers and infrastructure engineers to shape scalable tooling and compute stacks.

The role blends high-impact hands-on work with systems design, offering a competitive package and a dynamic research-driven

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

NLP Performance Engineer
NLP Performance Engineer

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Competitive compensation + bonus
Lunch provided
35 days annual leave
+5
LLM Performance Engineer — Scale Frontiers in AI
LLM Performance Engineer — Scale Frontiers in AI

Isomorphic Labs • Greater London

Hybrid
GBP 80,000 - 120,000
LLM Pretraining Scaling Engineer
LLM Pretraining Scaling Engineer

Humanloop • Greater London

On-site
GBP 57,000 - 73,000
Visa sponsorship where possible
Generous vacation & parental leave
Equity donation matching
ML Performance Engineer: Scale & Optimize GPU/CPU Workloads
ML Performance Engineer: Scale & Optimize GPU/CPU Workloads

G-Research • Greater London

On-site
GBP 90,000 - 135,000
Lunch provided (Just Eat for Business)
Barista bar
35 days annual leave
+5
Production AI Engineer: LLMs, NLP & Market Intelligence
Production AI Engineer: LLMs, NLP & Market Intelligence

Permutable.AI • Greater London

Hybrid
GBP 90,000 - 130,000
NLP & LLM Engineer — Search & Retrieval at Scale (Hybrid)
NLP & LLM Engineer — Search & Retrieval at Scale (Hybrid)

Opus Recruitment Solutions • Greater London

Hybrid
GBP 110,000
25 days holiday
Hybrid working model
Support for ongoing development
Senior LLM Serving Platform Engineer
Senior LLM Serving Platform Engineer

Scale AI • Greater London

On-site
GBP 90,000 - 130,000
Researcher, Training - London
Researcher, Training - London

United States Digital Space LLC • Greater London

Hybrid
GBP 70,000 - 90,000
Relocation support
Hybrid work schedule
LLM Researcher: Architect & Optimizer (London, Hybrid)
LLM Researcher: Architect & Optimizer (London, Hybrid)

OpenAI • Greater London

Hybrid
GBP 80,000 - 110,000
Relocation assistance
Hybrid work model
AI Infra Engineer for Scalable LLM Serving
AI Infra Engineer for Scalable LLM Serving

Scale AI, Inc. • Greater London

On-site
GBP 110,000 - 160,000