Platform Observability & Automation Engineer

Radiant

Greater London

On-site

GBP 90,000 - 120,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Radiant is seeking an Infrastructure Tooling & Observability Engineer to enhance reliability across its global GPU-focused infrastructure. You will design internal control planes, convert high-volume telemetry into actionable insights, and drive automation to reduce toil and improve system resilience.

Joining a fast-growing, engineering-led team, you will work with SRE and infrastructure engineers to deliver scalable tooling, CI/CD automation, and robust observability across distributed

Qualifications

  • Degree in Computer Science/Software Engineering, or equivalent experience.
  • 6–8 years of experience in infrastructure engineering, DevOps, SRE, and/or software engineering roles with a focus on operational systems.
  • Proven experience in at least one DevOps or software engineering role, building or maintaining production infrastructure tooling or platform systems.
  • Experience working in large-scale or distributed infrastructure environments.
  • Strong programming ability in Ruby (Rails) or Go, with willingness to work across multiple languages.
  • Hands-on experience with Ansible and AWX.
  • Strong experience with observability systems, including Grafana/Prometheus/Loki/Mimir/Grafana Alloy.
  • Familiarity with SNMP and syslog.
  • Experience working with Kubernetes in production environments.
  • Understanding of API design and REST-based services and service-to-service communication.
  • Experience building and maintaining CI/CD pipelines, including GitHub Actions and self-hosted runners.
  • Strong understanding of operational reliability concepts, including monitoring, alerting, capacity planning, and incident response.
  • Comfortable working closely with SRE, Platform Engineering, and infrastructure teams to translate operational needs into maintainable software systems.

Responsibilities

  • Design, build, and evolve internal tooling and observability platforms that support large-scale infrastructure operations across distributed environments.
  • Develop systems that turn high-volume telemetry (logs, metrics, events) into actionable insight, improving visibility, alerting quality, and operational decision-making.
  • Translate SRE reliability requirements into scalable, production-ready software solutions, including automation for incident detection, prevention, and remediation.
  • Drive automation across infrastructure operations, reducing manual effort in areas such as environment provisioning, cluster onboarding, inventory management, and lifecycle workflows.
  • Build tooling for capacity management, performance testing, benchmarking, and automated collection and analysis of results.
  • Contribute to Continual Service Improvement (CSI) initiatives by identifying operational inefficiencies and delivering durable engineering solutions.
  • Work closely with SRE and infrastructure engineering teams to embed observability and reliability into core platform workflows.
  • Interface with Platform Engineering teams to ensure tooling aligns with broader orchestration and infrastructure strategy.
  • Integrate and extend existing systems written in Ruby/Rails and Go, contributing to a consistent and maintainable engineering ecosystem.
  • Develop and maintain automation workflows using Ansible and AWX.
  • Support CI/CD-driven operational tooling, including GitHub Actions and self-hosted runners.

Skills

Programming (Ruby/Go)
DevOps/SRE
Kubernetes experience
CI/CD pipelines
Observability concepts

Education

Bachelor's degree in CS/Software Engineering

Tools

Ansible
AWX
Grafana
Prometheus
Loki
GitHub Actions
Self-hosted runners
REST APIs

Job description

Radiant is seeking an Infrastructure Tooling & Observability Engineer to enhance reliability across its global GPU-focused infrastructure. You will design internal control planes, convert high-volume telemetry into actionable insights, and drive automation to reduce toil and improve system resilience.

Joining a fast-growing, engineering-led team, you will work with SRE and infrastructure engineers to deliver scalable tooling, CI/CD automation, and robust observability across distributed

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Tooling & Observability Engineer( UK)
Infrastructure Tooling & Observability Engineer( UK)

Radiant • Greater London

On-site
GBP 90,000 - 120,000
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Staff Observability Engineer: Profiling & Telemetry
Staff Observability Engineer: Profiling & Telemetry

AI Startups UK • Greater London

Hybrid
GBP 325,000 - 390,000
Office space
Flexible working hours
Generous vacation
+2
Staff Observability Engineer — Fleet Telemetry & Profiling
Staff Observability Engineer — Fleet Telemetry & Profiling

Jackalope Digital LLC • Greater London

Hybrid
GBP 325,000 - 390,000
Visa sponsorship available
Senior Backend Engineer — AI Infra & Kubernetes Leader
Senior Backend Engineer — AI Infra & Kubernetes Leader

Radiant • Greater London

On-site
GBP 110,000 - 150,000
25 days leave
Private medical insurance
Cycle to Work
+3
24/7 Cloud Infra Support Engineer for AI & HPC
24/7 Cloud Infra Support Engineer for AI & HPC

Radiant • Greater London

On-site
GBP 52,000 - 80,000
25 days annual leave
Private medical insurance (Bupa)
Cycle to Work Scheme
+2
Cluster Architect
Cluster Architect

Radiant • Greater London

On-site
GBP 120,000 - 190,000
25 days leave
Medical insurance
Cycle to Work
+3
Staff Software Engineer, Observability & Profiling
Staff Software Engineer, Observability & Profiling

United States Digital Space LLC • Greater London

On-site
GBP 100,000 - 140,000
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability
Senior HPC/AI Infra SRE — 24/7 GPU Compute Reliability

Radiant • England

On-site
GBP 70,000 - 90,000
Exposure to industry-leading GPU and AI infrastructure
Collaborative, inclusive, and supportive engineering culture
Real ownership and influence over operational excellence
Cloud Infrastructure Support Engineer
Cloud Infrastructure Support Engineer

Radiant • Greater London

On-site
GBP 52,000 - 80,000
25 days annual leave
Private medical insurance (Bupa)
Cycle to Work Scheme
+2