ML System Engineer - (Reinforcement Learning)

Roc Search

Greater London

Hybrid

GBP 70,000 - 110,000

Full time

19 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Roc Search is assisting a London-based AI startup in finding an Infrastructure Engineer to design and scale the deployment layer of a distributed platform. You will orchestrate complex workloads across on-premise hardware, ensuring reliable execution within memory and compute constraints while collaborating with research and engineering teams.

Key duties include building testing environments, scaling pipelines, and turning experimental concepts into production software, with a focus on

Qualifications

  • Production-level Python development with clean testing and modular design.
  • Experience with distributed computing and state synchronization.
  • Memory and compute resource management across local tasks.
  • Telemetry, logging and alerting in distributed or containerized environments.

Responsibilities

  • Design runtime isolation and task scheduling for multiple local processes.
  • Build robust reconciliation mechanisms for atomic updates across remote environments.
  • Architect deployment flows with progressive rollout, validation, safety guardrails.
  • Develop large-scale virtualized environments to emulate real networks for validation.
  • Construct fault-tolerant distributed processing networks with automated recovery.
  • Profile and optimize performance; tune resources and hardware acceleration.

Skills

Python
MLops
Pytorch
Grafana
Elastic Search
TensorRT
Distributed computing
Systems engineering
Resource management
On-premise hardware

Job description

London (1-2 days per week in office)

AI Startup

Skills: Python, MLops, Pytorch, Grafana, Elastic Search, TensorRT

About the Company

Our client is a venture-backed technology company building software to optimise large-scale physical infrastructure. Their platform processes data locally at the source using decentralized networks. Their mission is to make global industrial operations more resilient, secure, and sustainable through advanced automation.

About the Role

Our client is seeking an Infrastructure Engineer to design and scale the deployment layer of their distributed technology platform. The core challenge involves orchestrating complex, concurrent software applications and analytical workloads across a massive network of diverse, on-premise hardware installations.

In this position, you will own the systems engineering required to guarantee that these disparate applications execute reliably within strict memory and compute constraints. You will also collaborate directly with their research and engineering teams, building robust testing environments, scaling distributed pipelines, and converting experimental concepts into dependable production software.

Key Responsibilities
Network Orchestration & System Performance
  • Resource Management: Design runtime isolation, task scheduling, and resource allocations for multiple concurrent local processes sharing the same hardware.
  • System Synchronization: Build robust reconciliation mechanisms to ensure atomic updates and version alignment across remote environments.
  • Release Management: Architect deployment flows supporting progressive rollout strategies, passive validation modes, safety guardrails, and automated recovery loops.
Scalable Infrastructure & Automation
  • Simulation Frameworks: Develop and maintain large-scale virtualized environments to safely emulate real-world networks and system behaviors for validation.
  • Pipeline Automation: Construct fault-tolerant distributed processing networks that support automated state saving, failure recovery, and cross-site data flows.
  • Performance Optimization: Profile system execution to improve processing performance through code optimization, resource tuning, and hardware acceleration on varied architectures.
Diagnostics & Engineering Standards
  • Data Streams: Establish reliable telemetry and ingestion channels that preserve data lineage for downstream analytic workflows.
  • System Telemetry: Implement comprehensive dashboard metrics, unified logging, and warning systems to detect system degradation or variance early.
  • Engineering Rigour: Troubleshoot deep architectural bugs, assist engineering teams with technical blockers, and enforce high coding standards via thorough review processes.
Candidate Profile
Core Technical Experience
  • Software Foundations: Exceptional software engineering capabilities in production-level Python, with a strong focus on clean testing patterns and modular design.
  • Distributed Computing: Extensive experience managing state alignment, messaging, and system execution across inconsistent networks and varied hardware form factors.
  • Resource Partitioning: Demonstrated skill in managing system memory, computing bounds, and storage across competing local application tasks.
  • Systems Infrastructure: Deep operational familiarity with managing background workloads, handling checkpointing/recovery, and optimizing software performance.
  • Production Operations: Solid track record establishing telemetry, tracking system health, and managing alerting rules in distributed or containerized ecosystems.
Preferred Technical Exposure
  • Experience building infrastructure for simulation software, virtual testbeds, or automated control systems.
  • Experience with Reinforcement Learning tools such as Ray RLlib
  • Familiarity with high-efficiency runtime environments or specialized hardware acceleration toolchains.
  • Background in remote system provisioning, telemetry transport protocols (such as messaging queues), or remote software updates.
  • Experience maintaining custom hardware environments, private network setups, or software for highly regulated environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied ML Scientist
Applied ML Scientist

Oliver Bernard • Greater London

Hybrid
GBP 80,000 - 110,000
Senior ML Engineer
Senior ML Engineer

C4S Search Ltd • Greater London

Hybrid
GBP 90,000 - 150,000
ML Systems Engineer
ML Systems Engineer

IC Resources • City of Edinburgh

Hybrid
GBP 65,000 - 100,000
Principal Machine Learning Engineer – Production Systems
Principal Machine Learning Engineer – Production Systems

SoftInWay Inc • Bristol

On-site
GBP 70,000 - 100,000
Machine Learning Engineer
Machine Learning Engineer

Searchability® • Greater London

Hybrid
GBP 90,000 - 120,000
ML Ops Engineer
ML Ops Engineer

cmcmarkets • Greater London

On-site
GBP 90,000 - 130,000
AI/ ML Infrastructure Engineer
AI/ ML Infrastructure Engineer

OpenSourced - Search & Selection • Bristol

Hybrid
GBP 90,000 - 110,000
Work on real-world AI systems
Direct impact on robotics capability
Fast-moving engineering environment
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Xcede • Greater London

Hybrid
GBP 120,000 - 150,000
ML Systems Engineer
ML Systems Engineer

Intellectual Capital Resources • City of Edinburgh

Hybrid
GBP 59,000 - 99,000
Principal ML Engineer
Principal ML Engineer

Anson McCade • Greater London

Hybrid
GBP 120,000 - 180,000
Annual Bonus
Private Medical Insurance
Enhanced Pension
+3