ML Systems Engineer (RL)

Roc Search Inc.

Greater London

Hybrid

GBP 100.000 - 110.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Mach aus dieser Rolle ein Bewerbungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

AI Startup in London is seeking a Machine Learning Systems Engineer to design and scale the deployment layer of a distributed platform. You will orchestrate concurrent apps across a vast on-premise network, ensuring reliable execution within memory and compute constraints.

You will collaborate with research and engineering teams, build testing environments, and scale distributed pipelines while turning experimental concepts into production software. Hybrid in-office pattern with weekly schedule.

Qualifikationen

  • Proficient production-grade Python development with strong testing patterns.
  • Experience delivering distributed systems and concurrent workloads.
  • Familiarity with monitoring and observability tooling in distributed apps.
  • Knowledge of memory and compute constraints in large-scale deployments.
  • Experience with ML platforms and model deployment workflows.

Aufgaben

  • Resource Management: design runtime isolation, scheduling, and allocations for multiple local processes.
  • System Synchronization: build reconciliation to ensure atomic updates and version alignment across environments.
  • Release Management: architect deployment flows with progressive rollout, validation modes, guardrails, and recovery loops.
  • Simulation and testing: create robust testing environments for distributed workloads.
  • Pipeline Automation: build fault-tolerant networks with state saving and cross-site data flows.
  • Performance optimization: profile and tune code and hardware acceleration.
  • Telemetry: establish data streams and dashboards for system health.
  • Engineering Rigour: assist teams with blockers and enforce coding standards.
  • Documentation and knowledge sharing for robust operations.

Kenntnisse

Python
MLops
Pytorch
Grafana
Elastic Search
TensorRT

Jobbeschreibung

Machine Learning Systems Engineer (Distributed Systems)
London (1-2 days per week in office)
AI Startup
£100,000-£110,000 DOE + Bonus
Skills: Python, MLops, Pytorch, Grafana, Elastic Search, TensorRT
About the Company

Our client is a venture-backed technology company building software to optimise large-scale physical infrastructure. Their platform processes data locally at the source using decentralized networks. Their mission is to make global industrial operations more resilient, secure, and sustainable through advanced automation.

About the Role

Our client is seeking an Infrastructure Engineer to design and scale the deployment layer of their distributed technology platform. The core challenge involves orchestrating complex, concurrent software applications and analytical workloads across a massive network of diverse, on-premise hardware installations.

In this position, you will own the systems engineering required to guarantee that these disparate applications execute reliably within strict memory and compute constraints. You will also collaborate directly with their research and engineering teams, building robust testing environments, scaling distributed pipelines, and converting experimental concepts into dependable production software.

Key Responsibilities
Network Orchestration & System Performance
  • Resource Management: Design runtime isolation, task scheduling, and resource allocations for multiple concurrent local processes sharing the same hardware.
  • System Synchronization: Build robust reconciliation mechanisms to ensure atomic updates and version alignment across remote environments.
  • Release Management: Architect deployment flows supporting progressive rollout strategies, passive validation modes, safety guardrails, and automated recovery loops.
Scalable Infrastructure & Automation
  • Simulation Frameworks: Develop and maintain large-scale virtualized environments to safely emulate real-world networks and system behaviors for validation.
  • Pipeline Automation: Construct fault-tolerant distributed processing networks that support automated state saving, failure recovery, and cross-site data flows.
  • Performance Optimization: Profile system execution to improve processing performance through code optimization, resource tuning, and hardware acceleration on varied architectures.
Diagnostics & Engineering Standards
  • Data Streams: Establish reliable telemetry and ingestion channels that preserve data lineage for downstream analytic workflows.
  • System Telemetry: Implement comprehensive dashboard metrics, unified logging, and warning systems to detect system degradation or variance early.
  • Engineering Rigour: Troubleshoot deep architectural bugs, assist engineering teams with technical blockers, and enforce high coding standards via thorough review processes.
Candidate Profile
Core Technical Experience
  • Software Foundations: Exceptional software engineering capabilities in production-level Python, with a strong focus on clean testing patterns and modular design.
  • Distributed Computing: Extensive experience managing state alignment, messaging, and system execution across inconsistent networks and varied hardware form factors.
  • Resource Partitioning: Demonstrated skill in managing system memory, computing bounds, and storage across competing local application tasks.
  • Systems Infrastructure: Deep operational familiarity with managing background workloads, handling checkpointing/recovery, and optimizing software performance.
  • Production Operations: Solid track record establishing telemetry, tracking system health, and managing alerting rules in distributed or containerized ecosystems.
Preferred Technical Exposure
  • Experience building infrastructure for simulation software, virtual testbeds, or automated control systems.
  • Experience with Reinforcement Learning tools such as Ray RLlib
  • Familiarity with high-efficiency runtime environments or specialized hardware acceleration toolchains.
  • Background in remote system provisioning, telemetry transport protocols (such as messaging queues), or remote software updates.
  • Experience maintaining custom hardware environments, private network setups, or software for highly regulated environments.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Machine Learning Engineer
Machine Learning Engineer

Platform Recruitment • Greater London

Vor Ort
GBP 42.000 - 70.000
Senior ML Engineer
Senior ML Engineer

C4S Search Ltd • Greater London

Vor Ort
GBP 90.000 - 150.000
Senior ML Ops Engineer 201043
Senior ML Ops Engineer 201043

Harnham • Großbritannien

Vor Ort
GBP 70.000 - 110.000
Ongoing training
Ownership of ML Ops function
AI/ ML Infrastructure Engineer
AI/ ML Infrastructure Engineer

OpenSourced - Search & Selection • Bristol

Vor Ort
GBP 99.000 - 121.000
Work on real-world AI systems
Direct impact on robotics capability
Fast-moving engineering environment
Principal Machine Learning Engineer – Production Systems
Principal Machine Learning Engineer – Production Systems

SoftInWay Inc • Bristol

Vor Ort
GBP 70.000 - 100.000
Distributed ML Systems Engineer - Hybrid (London)
Distributed ML Systems Engineer - Hybrid (London)

Roc Search Inc. • Greater London

Hybrid
GBP 100.000 - 110.000
ML Platform Engineer
ML Platform Engineer

Xist4 IT Limited. • City Of London

Remote
GBP 90.000 - 115.000
Senior MLOps Engineer - AI Infrastructure
Senior MLOps Engineer - AI Infrastructure

Harnham • Greater London

Hybrid
GBP 51.000 - 85.000
Hybrid work model
MLOps Engineer
MLOps Engineer

Harnham - Data & Analytics Recruitment • Greater London

Hybrid
GBP 75.000 - 85.000
MLOps Engineer
MLOps Engineer

Harnham • Greater London

Hybrid
GBP 75.000 - 85.000
Bonus up to 10%
Private healthcare
Hybrid work London