Remote Lead SRE: Scale Observability & Data Infrastructure

Randstad Technologies Recruitment

Ribble Valley

On-site

GBP 95,000 - 135,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Randstad Technologies Recruitment in the United Kingdom is seeking a Lead Site Reliability Engineer (SRE) focused on observability and reliability. You will design, scale, and operate massive observability systems, joining an autonomous team of engineers dedicated to keeping our global services online and performant.

Remote role with potential travel; you will own complex data infrastructure challenges and drive best practices in monitoring, alerting, and incident response for large-scale

Qualifications

  • 5+ years operating mid-to-large distributed systems on Linux VMs or bare-metal.
  • 2+ years developing in Go, Python, Ruby, Scala, or Bash.
  • Hands-on experience with Prometheus/Thanos/Cortex, Kafka, the ELK stack, Ansible, or Consul.
  • Comfortable diving into unfamiliar codebases and participating in an on-call rotation.

Responsibilities

  • Scale Prometheus metrics infrastructure to handle 100+ million active series.
  • Operate large Elasticsearch clusters holding 2000+TB of data.
  • Grow high-throughput Kafka data pipelines processing hundreds of thousands of events per second.
  • Build custom alerting workflows and self-service APIs for internal engineering teams.
  • Provision cloud and private infrastructure using Terraform.

Skills

Distributed systems
Linux
On-call rotation
Go/Python/Ruby/Scala/Bash

Tools

Prometheus/Thanos/Cortex
Kafka
ELK stack
Ansible
Consul
Terraform

Job description

Randstad Technologies Recruitment in the United Kingdom is seeking a Lead Site Reliability Engineer (SRE) focused on observability and reliability. You will design, scale, and operate massive observability systems, joining an autonomous team of engineers dedicated to keeping our global services online and performant.

Remote role with potential travel; you will own complex data infrastructure challenges and drive best practices in monitoring, alerting, and incident response for large-scale

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000
SRE Architect: Reliability Leader, Observability
SRE Architect: Reliability Leader, Observability

Hitachi Automotive Systems Americas, Inc. • Greater London

On-site
GBP 95,000 - 130,000
Senior SRE: Cloud Reliability & Observability Lead
Senior SRE: Cloud Reliability & Observability Lead

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Senior Site Reliability Engineer - Scale & Observability Lead
Senior Site Reliability Engineer - Scale & Observability Lead

Tata Consultancy Services • London

On-site
GBP 60,000 - 80,000
Lead SRE: Scale, Reliability & Automation
Lead SRE: Scale, Reliability & Automation

McLaren Automotive Ltd • Woking

On-site
GBP 70,000 - 90,000
SRE Manager: Scale, Observability & Cloud Resilience
SRE Manager: Scale, Observability & Cloud Resilience

Holland Barrett • Greater London

Hybrid
GBP 90,000 - 120,000
33 days holiday
Private medical insurance
Annual bonus up to 10%
+5
Senior Lead SRE - Reliability & Observability Leader
Senior Lead SRE - Reliability & Observability Leader

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 130,000
Senior AWS SRE Lead — Cloud Reliability & Automation
Senior AWS SRE Lead — Cloud Reliability & Automation

United States Digital Space LLC • Greater London

Hybrid
GBP 95,000 - 130,000
Hybrid working
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
SRE Leadership: Reliability, Automation & Observability
SRE Leadership: Reliability, Automation & Observability

Jobtailor • Bristol

On-site
GBP 110,000 - 140,000