Founding Inference Research Engineer

Gradiant

Greater London

Hybrid

GBP 110,000 - 140,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Gradiant is seeking an engineer to identify bottlenecks across the inference stack for open-weight models. You will measure targets on NVIDIA and AMD hardware, design kernels and memory layouts, and push changes to runtimes and serving paths.

The role emphasizes turning research into production services and sharing findings publicly. You should be proficient in Python and a systems language (C++ or Rust), with experience in GPU programming or low-level hardware considerations and a strong grasp

Qualifications

  • Experience building ML systems, compilers, runtimes, or kernels.
  • Ability to explain decisions and results in detail.
  • Proficiency in Python and a systems language (C++ or Rust).
  • CUDA or HIP kernel writing or strong low-level GPU experience.

Responsibilities

  • Measure models and workloads on NVIDIA and AMD hardware.
  • Design and implement kernels, formats, memory layouts, and cache policies.
  • Modify runtimes, schedulers, and serving paths to improve performance.
  • Test latency, throughput, energy use, and model correctness on real workloads.
  • Turn research into production services.
  • Share research publicly to educate others.
  • Collaborate with silicon team to identify bottlenecks and opportunities.

Skills

ML systems
Python
C++ or Rust
CUDA/HIP kernels
Profiling
Transformer execution
Production code

Tools

CUDA
GPU kernels

Job description

Find and remove bottlenecks across model representation, GPU kernels, runtimes, and serving for open-weight models.

Why this role exists

We are building an inference research team to find where current hardware and software waste time, memory, and energy when serving open-weight models.

You will measure those bottlenecks, test changes across the inference stack, and turn useful results into production systems. Your findings will also inform our silicon work.

What you will do
  • Measure target models and workloads on current NVIDIA and AMD hardware.
  • Design and implement kernels, numerical formats, memory layouts, and cache policies.
  • Change runtimes, schedulers, distributed execution, and serving paths when they limit the system.
  • Test latency, throughput, energy use, model quality, correctness, and reliability on real workloads.
  • Turn research results into production services.
  • Share research publicly so that others can learn about how inference research is done.
  • Work with the silicon team to identify hardware bottlenecks and design opportunities.
Problems you may work on
  • Kernels and fused execution paths for exact model shapes.
  • Low-bit formats that preserve model quality.
  • KV-cache layout, movement, compression, and reuse.
  • Sparse and mixture-of-experts execution.
  • Speculative and parallel decoding.
  • Prefill and decode scheduling across multiple GPUs.
  • Performance models that explain the gap between hardware limits and measured results.
  • Runtime designs for long-running agent workloads.
What we are looking for
  • You have built an ML systems, compiler, runtime, or kernel project, and you can explain your decisions and results in detail.
  • You can work in Python and a systems language such as C++ or Rust. We use C++ mostly.
  • You can write CUDA or HIP kernels, or you have strong low-level systems experience and can learn GPU programming.
  • You understand transformer execution, including attention, memory movement, and numerical precision.
  • You use profiling and controlled experiments to identify system limits.
  • You are open to learning about hardware, and what is limiting the execution of the model.
  • You can maintain production code, not only experimental code.
  • You can explain hard technical solutions clearly
This role is not for you if
  • You want to work within one layer of the inference stack.
  • You want framework integration to be the main technical work.
  • You want papers or notebook results to be the primary output.
  • You aren't excited by rewriting things from scratch, and taking time to find the globally optimal solution to a problem.
  • You need a complete specification before you can investigate a problem.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Silicon Engineer
Founding Silicon Engineer

Gradiant • United Kingdom

Remote
GBP 70,000 - 110,000
Open-Weight Inference Systems Architect
Open-Weight Inference Systems Architect

Gradiant • Greater London

Hybrid
GBP 110,000 - 140,000
Open-Weight Inference Systems Architect
Open-Weight Inference Systems Architect

Gradiant • Greater London

Hybrid
GBP 110,000 - 140,000
ML Inference Engineer
ML Inference Engineer

United States Digital Space LLC • Greater London

On-site
GBP 80,000 - 120,000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Inference
Inference

Genesis AI • Greater London

On-site
GBP 110,000 - 150,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 80,000 - 120,000
Equity
Head of Silicon
Head of Silicon

Gradiant • United Kingdom

Remote
GBP 120,000 - 190,000
AI Inference Engineer
AI Inference Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity
Biannual bonus
Fully expensed tech
+2
Senior Software Engineer (Inference)
Senior Software Engineer (Inference)

AssemblyAI • Greater London

On-site
GBP 90,000 - 150,000
Home office stipend
Equity grant
Premium medical, dental, vision plans
+4