Lead Site Reliability Engineer

Factset

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health coverage
Free lunch in the office (Mon–Fri)
Employee social events and sports

Job summary

Factset is seeking a Lead Site Reliability Engineer to ensure reliability, scalability, and performance of critical systems. You will collaborate with development and operations to automate processes and embed reliability into services from design through deployment.

The role involves defining SLOs/SLIs, leading incident response, and driving continuous improvements across teams in a hybrid working model in the Greater London area.

Qualifications

  • Must have hands-on Kubernetes deployment and operations experience.
  • Experience with Helm for packaging and deploying apps.
  • Strong troubleshooting and incident‑response abilities.

Responsibilities

  • Ensure reliability, scalability, and performance of systems and services.
  • Respond to incidents, perform post-mortems, and drive improvements.
  • Define and track SLOs/SLIs and collaborate with teams to build in reliability.
  • Automate processes, reduce toil, and improve operational efficiency.
  • Participate in on‑call rotations and assist capacity planning.

Skills

Kubernetes
Helm
English fluency
Networking
Storage
Security
Open-source
SRE principles
DevOps
Incident response

Education

Bachelors degree in Computer Science or related field

Job description

  • We are looking for a skilled and motivated Lead Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our systems and services. You will work closely with development and operations teams to build and maintain robust infrastructure, automate processes, and drive engineering best practices
  • Monitor, maintain, and improve the reliability and availability of production systems
  • Respond to and resolve incidents, conducting thorough post-mortems to prevent recurrence
  • Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Collaborate with development teams to build reliability into services from the ground up
  • Design and implement automation to reduce toil and improve operational efficiency
  • Participate in an on-call rotation to support critical systems
  • Contribute to capacity planning and performance optimization efforts
  • Document systems, processes, and runbooks to support the wider team
Benefits
  • Comprehensive health coverage for employees and their families, at little or no cost to employees
  • Free working lunch in the office Monday through Friday
  • A social community involved in sports, charities, and in-office events
  • Certification reimbursement for eligible expenses related to the CFA, IPM, CAIA, and FRM exams
  • Generous PTO for personal, vacation, parental and medical leave
  • Wellness discounts across gyms and other facilities

Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and IngressHands-on experience deploying, managing, and troubleshooting workloads in KubernetesMust be fluent in English both verbal and writtenUnderstanding of Kubernetes networking, storage, and security best practicesFamiliarity with Helm for application packaging and deploymentBachelors degree in computer science or relevant degreeExperience with Kubernetes cluster management and administrationWilling to work a hybrid modelCommitment to a blameless culture and continuous learningStrong problem-solving and analytical skills with a methodical approach to troubleshootingExcellent communication skills with the ability to collaborate across technical and non-technical teamsAbility to work effectively under pressure, particularly during incident responseA proactive mindset with a focus on automation and continuous improvementExperience contributing to open-source projectsFamiliarity with SRE principles as defined by the Google SRE handbookPrevious experience in a DevOps or Platform Engineering role

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

iXceed Solutions • Basildon

On-site
GBP 55,000 - 75,000
SRE Engineer
SRE Engineer

Selby Jennings • Greater London

On-site
GBP 90,000 - 130,000
Senior Site Reliability Engineer (Platform Reliability, Resilience)
Senior Site Reliability Engineer (Platform Reliability, Resilience)

Elastic • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage for you and family
Flexible location & schedule
Generous vacation days
+3
Lead Site Reliability Engineer | Kubernetes & Automation
Lead Site Reliability Engineer | Kubernetes & Automation

Factset • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage
Free lunch in the office (Mon–Fri)
Employee social events and sports
Platform Reliability Engineer — IaC, Kubernetes & GCP
Platform Reliability Engineer — IaC, Kubernetes & GCP

Selby Jennings • Greater London

On-site
GBP 90,000 - 130,000
Backend Engineer
Backend Engineer

TESTQ Technologies LTD. • Sheffield

Hybrid
GBP 70,000 - 110,000
Senior Site Reliability Engineer - Kubernetes & Automation
Senior Site Reliability Engineer - Kubernetes & Automation

20035 FactSet Europe Limited • Greater London

Hybrid
GBP 90,000 - 150,000
Site Reliability Engineer (SRE) - Cloud Kubernetes Platform
Site Reliability Engineer (SRE) - Cloud Kubernetes Platform

Intuition IT Solutions Ltd • Glasgow

Hybrid
GBP 60,000 - 75,000
Hybrid work model
Lead SRE for Kubernetes Reliability & Automation
Lead SRE for Kubernetes Reliability & Automation

Factset • Greater London

Hybrid
GBP 90,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000