Senior Inference Architect — Low-Latency, On-Device & GPU

Genesis AI

Greater London

On-site

GBP 110,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Genesis AI is seeking an expert in distributed systems and ML infrastructure to build and optimize low-latency inference pipelines for on-device robotics. The role focuses on GPU-accelerated architectures and integrating high-performance kernels into user-friendly frameworks.

Candidates should demonstrate production-grade Python skills with a strong background in C++/Rust/Go, and a track record of scaling workloads across both on-device and cluster environments.

Qualifications

  • 8+ years in distributed systems, ML infra, or high-performance serving.
  • Production-grade Python with strong systems language background (C++, Rust, Go).
  • Expertise in low-level optimization: CUDA, Triton, kernels, memory scheduling.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment in robotics.
  • Design and optimize distributed inference systems on GPU clusters for high-throughput and efficiency.
  • Implement efficient low-level code (CUDA, Triton) and integrate into high-level frameworks.
  • Optimize workloads for throughput and latency via batching, scheduling, and memory management.
  • Develop monitoring and debugging tools to ensure reliability and rapid regression diagnosis.

Skills

Distributed systems
Python
C++/Rust/Go
Low-level performance
ML infrastructure
GPU inference

Tools

CUDA
Triton

Job description

Genesis AI is seeking an expert in distributed systems and ML infrastructure to build and optimize low-latency inference pipelines for on-device robotics. The role focuses on GPU-accelerated architectures and integrating high-performance kernels into user-friendly frameworks.

Candidates should demonstrate production-grade Python skills with a strong background in C++/Rust/Go, and a track record of scaling workloads across both on-device and cluster environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference
Inference

Genesis AI • Greater London

On-site
GBP 110,000 - 150,000
Senior AI Training Infra Engineer - Distributed PyTorch
Senior AI Training Infra Engineer - Distributed PyTorch

Genesis AI • Greater London

Hybrid
GBP 120,000 - 170,000
Senior Real-Time ML Inference Engineer
Senior Real-Time ML Inference Engineer

BITKRAFT Ventures • United Kingdom

On-site
GBP 140,000 - 200,000
Lead AI Training Infrastructure Engineer
Lead AI Training Infrastructure Engineer

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)
AI Inference Engineer — GPU-Optimized Rust/Python (Equity)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
Senior ML Engineer: GPU Inference & Low-Precision Training
Senior ML Engineer: GPU Inference & Low-Precision Training

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior DL Inference Engineer — GPU-Accelerated AI at Scale
Senior DL Inference Engineer — GPU-Accelerated AI at Scale

NVIDIA • United Kingdom

On-site
GBP 110,000 - 160,000
Competitive salaries
Extensive benefits package
Diversity & inclusion
Senior GPU Systems Engineer: Large-Scale Inference & RL
Senior GPU Systems Engineer: Large-Scale Inference & RL

Reflection • Greater London

On-site
GBP 70,000 - 100,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
GPU Performance Engineer: Scale Inference & Training
GPU Performance Engineer: Scale Inference & Training

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
AI Inference Engineer (Staff) - GPU ML Systems + Equity
AI Inference Engineer (Staff) - GPU ML Systems + Equity

Perplexity • Greater London

On-site
GBP 80,000 - 120,000
Equity