Inference

Genesis AI

Greater London

On-site

GBP 110,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Genesis AI is seeking an expert in distributed systems and ML infrastructure to build and optimize low-latency inference pipelines for on-device robotics. The role focuses on GPU-accelerated architectures and integrating high-performance kernels into user-friendly frameworks.

Candidates should demonstrate production-grade Python skills with a strong background in C++/Rust/Go, and a track record of scaling workloads across both on-device and cluster environments.

Qualifications

  • 8+ years in distributed systems, ML infra, or high-performance serving.
  • Production-grade Python with strong systems language background (C++, Rust, Go).
  • Expertise in low-level optimization: CUDA, Triton, kernels, memory scheduling.

Responsibilities

  • Build low-latency inference pipelines for on-device deployment in robotics.
  • Design and optimize distributed inference systems on GPU clusters for high-throughput and efficiency.
  • Implement efficient low-level code (CUDA, Triton) and integrate into high-level frameworks.
  • Optimize workloads for throughput and latency via batching, scheduling, and memory management.
  • Develop monitoring and debugging tools to ensure reliability and rapid regression diagnosis.

Skills

Distributed systems
Python
C++/Rust/Go
Low-level performance
ML infrastructure
GPU inference

Tools

CUDA
Triton

Job description

What You’ll Do
  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics

  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization

  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks

  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)

  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)

  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)

  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling

  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments

  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • Greater London

Hybrid
GBP 120,000 - 170,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • Greater London

Hybrid
GBP 120,000 - 170,000
Senior Inference Architect — Low-Latency, On-Device & GPU
Senior Inference Architect — Low-Latency, On-Device & GPU

Genesis AI • Greater London

On-site
GBP 110,000 - 150,000
Training / AI Infrastructure Engineering & Research London
Training / AI Infrastructure Engineering & Research London

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 80,000 - 120,000
Equity
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Senior Machine Learning Applications and Compiler Engineer, LPX
Senior Machine Learning Applications and Compiler Engineer, LPX

NVIDIA Corporation • Cambridge

Hybrid
GBP 95,000 - 135,000
Founding Inference Research Engineer
Founding Inference Research Engineer

Gradiant • Greater London

Hybrid
GBP 110,000 - 140,000