Site Reliability Engineer

Jobtailor

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Jobtailor is seeking a senior Site Reliability Engineer to tackle reliability, scalability, and efficiency challenges across SRE and development teams. You will build and run large-scale, distributed fault-tolerant systems that power the Genesis platform, optimize existing infrastructure, and cut toil through automation to improve uptime.

The role demands strong Python/Go skills, deep distributed systems design, and leadership across teams.

Qualifications

  • Proven ability designing and troubleshooting distributed systems.
  • Experience leading complex, large-scale technical projects.
  • Strong cloud expertise across platforms and services.

Responsibilities

  • Take on ambiguous reliability, scalability, and efficiency challenges and drive solutions across SRE and development teams.
  • Build and run large-scale, massively distributed, fault-tolerant systems supporting the Genesis platform.
  • Optimize existing systems and build infrastructure.
  • Eliminate toil through automation to improve uptime and rate of change.
  • Cultivate a culture of reliability throughout the organization.
  • Guide technical decisions balancing system health with product priorities.
  • Ensure long-term health, maintainability, and reliability of services.
  • Perform capacity planning and performance analysis.
  • Proactively prevent incidents.
  • Work across teams to build robust, reusable solutions.

Skills

Python
Go
Distributed Systems
Kubernetes
Cloud Functions
NALSD
Technical Leadership
Cross-Team Collaboration
Algorithms & Complexity
Mentoring/Leadership

Tools

Kubernetes
Cloud Functions

Job description

Take on ambiguous reliability, scalability, and efficiency challenges and drive solutions across SRE and development teams
Build and run large-scale, massively distributed, fault-tolerant systems supporting the Genesis platform
Optimize existing systems and build infrastructure
Eliminate toil through automation to improve uptime and rate of change
Cultivate a culture of reliability throughout the organization
Guide technical decisions balancing system health with product priorities
Ensure long-term health, maintainability, and reliability of services
Perform capacity planning and performance analysis
Proactively prevent incidents
Work across teams to build robust, reusable solutions

Requirements
  • Strong software engineering skills in Python, Go, or similar
  • Extensive experience designing, analyzing, and troubleshooting distributed systems
  • Deep expertise with cloud computing platforms, including Kubernetes and Cloud Functions
  • Expertise in Non-Abstract Large Systems Design (NALSD)
  • Experience leading complex, large-scale technical projects
  • Experience providing technical leadership across teams
  • Ability to apply coding, algorithms, and complexity analysis to solve ambiguous problems at scale with minimal disruption
  • Collaborative and intellectually curious mindset
  • Comfortable working across a wide variety of backgrounds and bringing cross-team perspective
  • Personal website or GitHub, LinkedIn, and resume fields are available in the application form; the website, LinkedIn, and resume are not required
Core Competencies

Demonstrates strong software engineering skills in Python and Go, with extensive experience in designing and troubleshooting distributed systems. Proven ability to lead complex technical projects and optimize large-scale, fault-tolerant systems while fostering a culture of reliability and collaboration.

Highest-signal resume keywords
  • Python Programming
  • Go Programming
  • Distributed Systems Design
  • Kubernetes Expertise
  • Technical Leadership
Hard Skills
  • Software Engineering
  • System Optimization
  • Capacity Planning
  • Performance Analysis
  • Automation
  • Algorithms
  • Complexity Analysis
  • Fault-Tolerant Systems
  • Non-Abstract Large Systems Design
  • Incident Prevention
Soft Skills
  • Collaborative Mindset
  • Intellectual Curiosity
Industry Keywords
  • Reliability
  • Scalability
  • Efficiency
  • Cross-Team Collaboration
Tools & Technologies
  • Cloud Computing Platforms
  • Cloud Functions
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Genesis AI • Greater London

Hybrid
GBP 90,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Daily catered lunches
Modern office environment
Tech talks and knowledge sharing
Lead Site Reliability Engineer - Glasgow
Lead Site Reliability Engineer - Glasgow

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

P2P • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Site Reliability Engineer
Site Reliability Engineer

iXceed Solutions • Basildon

On-site
GBP 55,000 - 75,000
Vice President - Site Reliability Engineering
Vice President - Site Reliability Engineering

Goldman Sachs • Birmingham

On-site
GBP 60,000 - 100,000
Infrastructure Engineer
Infrastructure Engineer

Jobtailor • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Platform Engineer
Senior Platform Engineer

Jobtailor • Greater London

Hybrid
GBP 90,000 - 120,000
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000