Remote Staff ML Efficiency Engineer — Scale & Optimize

Reddit, Inc.

Greater London

On-site

GBP 110,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Global Benefits
Family Planning
Mental Health Support
Pension with match
Private Medical/Dental
Bike to Work
Paid Parental Leave

Job summary

Reddit, Inc. is seeking a senior ML platform engineer to design and optimize scalable ML training, inference, and serving systems. You will build tooling, profiling, and benchmarking capabilities to improve developer productivity and reduce costs.

You will collaborate with ML researchers and product teams to identify bottlenecks, increase GPU utilization, and drive platform reliability as workloads grow. Remote-friendly UK role with strong compensation.

Qualifications

  • BS/MS/PhD in Computer Science or a related field.
  • 5+ years of software engineering experience.
  • Strong proficiency in Python.
  • Proficiency in at least one systems language (Go, C++, Rust, or Java).
  • Experience building distributed systems at scale.
  • Experience with machine learning infrastructure, training systems, or model serving platforms.
  • Deep understanding of performance engineering and systems optimization.

Responsibilities

  • Design and build systems that improve ML training and inference efficiency.
  • Develop tooling to help ML engineers debug, profile, and monitor performance.
  • Improve GPU and resource utilization via scheduling, caching, and optimization.
  • Partner with researchers and product teams to drive performance improvements.
  • Build benchmarking frameworks and performance dashboards.
  • Optimize distributed training infrastructure and data pipelines.
  • Lead cross-functional initiatives to boost productivity of Reddit ML engineers.
  • Drive technical strategy for ML platform scalability, reliability, and cost efficiency.

Skills

Python
Distributed systems
Performance engineering
Debugging/profiling
Go/C++/Rust/Java
ML infrastructure
CS fundamentals

Education

BS/MS/PhD in Computer Science

Tools

PyTorch Distributed
Ray
TensorFlow
Spark

Job description

Reddit, Inc. is seeking a senior ML platform engineer to design and optimize scalable ML training, inference, and serving systems. You will build tooling, profiling, and benchmarking capabilities to improve developer productivity and reduce costs.

You will collaborate with ML researchers and product teams to identify bottlenecks, increase GPU utilization, and drive platform reliability as workloads grow. Remote-friendly UK role with strong compensation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Ads ML Platform Engineer - Scale Data Infra
Remote Ads ML Platform Engineer - Scale Data Infra

Reddit, Inc. • United Kingdom

Remote
GBP 70,000 - 90,000
Global Benefit programs
Family Planning Support
Mental Health & Coaching Benefits
+4
Remote ML Infrastructure Engineer — Ads Platform
Remote ML Infrastructure Engineer — Ads Platform

Reddit, Inc. • Greater London

On-site
GBP 70,000 - 110,000
Private Medical and Dental Scheme
Group Personal Pension Scheme with EMP
Flexible Vacation
+2
ML Performance Engineer: Large-Scale GPU/CPU Optimization
ML Performance Engineer: Large-Scale GPU/CPU Optimization

gresearch • Greater London

On-site
GBP 90,000 - 130,000
Lunch provided
35 days annual leave
9% pension contributions
+3
Remote Staff ML Engineer — Lead Production‑Grade AI
Remote Staff ML Engineer — Lead Production‑Grade AI

Hexwired Recruitment Limited • Ribble Valley

On-site
GBP 100,000 - 200,000
ML Platform Engineer - Scale Training & Inference
ML Platform Engineer - Scale Training & Inference

Dex • Greater London

On-site
GBP 180,000 - 300,000
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

On-site
GBP 75,000 - 95,000
Co-working benefit
Daily lunch program
Regular community and social events
ML Performance Engineer: Scale GPU/CPU ML Workloads
ML Performance Engineer: Scale GPU/CPU ML Workloads

G-Research • Greater London

Hybrid
GBP 90,000 - 150,000
Competitive pay
Lunch provided
Annual leave 35d
+5
Staff ML Platform Engineer: Scalable AI Compute & HPC
Staff ML Platform Engineer: Scalable AI Compute & HPC

Chemify Ltd • Glasgow

On-site
GBP 90,000 - 120,000
Senior ML Runtime Engineer for Scalable Inference
Senior ML Runtime Engineer for Scalable Inference

Fractile • Bristol

Hybrid
GBP 70,000 - 90,000
Competitive salary and equity
Hybrid working
Visible and valued contributions
Senior ML Engineer - Recommender Systems (Remote)
Senior ML Engineer - Recommender Systems (Remote)

Embedded Shishya • United Kingdom

Remote
GBP 91,000 - 111,000
RSUs
Remote work
Global culture