LLM Inference Performance Engineer

G-Research

Greater London

On-site

GBP 90,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary bonus
35 days leave
Pension contributions
Healthcare and life assurance
Cycle-to-work

Job summary

G-Research in London is seeking an exceptional NLP Performance Engineer to own large-scale LLM inference performance within the NLP Engineering team. You’ll optimise the inference stack for cost-efficiency and speed, collaborating with researchers to bring ideas to life.

This hands-on role involves profiling workloads, improving deployment across GPU architectures, and building reliable tooling for ML workloads, while communicating across research, infra and engineering teams.

Qualifications

  • Bachelor’s, Master’s or PhD in computer science or equivalent experience.
  • Proven experience profiling, benchmarking and optimising large-scale LLM inference workloads.
  • Scientific, evidence-led approach to performance optimisation with rigorous benchmarking.

Responsibilities

  • Profiling, benchmarking and optimising large-scale LLM inference workloads across compute infrastructure.
  • Ensuring efficient deployment of latest models across GPU architectures; adapt inference stack as hardware evolves.
  • Designing and implementing inference optimisations while maintaining output quality.
  • Developing reference implementations, libraries and tooling to improve NLP workloads.
  • Collaborating with researchers, senior stakeholders and engineers to design optimised solutions.
  • Evolve the compute stack with systems, architecture and platform teams for long-term decisions.

Skills

LLM inferences
Performance benchmarking
Python
CUDA
PyTorch
Software engineering
Communication
Transformer inference

Education

Bachelor/Master/PhD in Computer Science

Tools

vLLM
SGLang
TensorRT-LLM
TGI
CUDA toolkit

Job description

G-Research in London is seeking an exceptional NLP Performance Engineer to own large-scale LLM inference performance within the NLP Engineering team. You’ll optimise the inference stack for cost-efficiency and speed, collaborating with researchers to bring ideas to life.

This hands-on role involves profiling workloads, improving deployment across GPU architectures, and building reliable tooling for ML workloads, while communicating across research, infra and engineering teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

NLP Performance Engineer: Scale LLM Inference for Research
NLP Performance Engineer: Scale LLM Inference for Research

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Competitive compensation + bonus
Lunch provided
35 days annual leave
+5
NLP Performance Engineer
NLP Performance Engineer

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Competitive compensation + bonus
Lunch provided
35 days annual leave
+5
Natural Language Programming Performance Engineer
Natural Language Programming Performance Engineer

G-Research • Greater London

On-site
GBP 90,000 - 150,000
Discretionary bonus
35 days leave
Pension contributions
+2
ML Performance Engineer: Large-Scale GPU/CPU Optimization
ML Performance Engineer: Large-Scale GPU/CPU Optimization

gresearch • Greater London

On-site
GBP 90,000 - 130,000
Lunch provided
35 days annual leave
9% pension contributions
+3
ML Inference & Serving Engineer
ML Inference & Serving Engineer

Google Inc. • Greater London

Hybrid
GBP 153,000 - 222,000
LLM Researcher: Architect & Optimizer (London, Hybrid)
LLM Researcher: Architect & Optimizer (London, Hybrid)

OpenAI • Greater London

Hybrid
GBP 80,000 - 110,000
Relocation assistance
Hybrid work model
Production AI Engineer: LLMs, NLP & Market Intelligence
Production AI Engineer: LLMs, NLP & Market Intelligence

Permutable.AI • Greater London

Hybrid
GBP 90,000 - 130,000
ML Performance Engineer: Scale GPU/CPU ML Workloads
ML Performance Engineer: Scale GPU/CPU ML Workloads

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
Senior LLM Serving Platform Engineer
Senior LLM Serving Platform Engineer

Scale AI • Greater London

On-site
GBP 90,000 - 130,000
NLP & LLM Engineer — Search & Retrieval at Scale (Hybrid)
NLP & LLM Engineer — Search & Retrieval at Scale (Hybrid)

Opus Recruitment Solutions • Greater London

Hybrid
GBP 110,000
25 days holiday
Hybrid working model
Support for ongoing development