Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

FactSet Research Systems Inc.

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

FactSet is seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, and performance of our systems in a hybrid UK environment. You will partner with development and operations teams to automate workflows, improve observability, and drive reliability initiatives across production services.

The role involves defining SLOs/SLIs, handling on-call rotations, and contributing to capacity planning while advocating for a blameless culture and continuous improvement.

Qualifications

  • Bachelor's degree in CS or related field.
  • Strong experience with Kubernetes and cloud-native tooling.
  • Proactive, analytical trouble-shooter with incident handling skills.

Responsibilities

  • Monitor, manage, and improve reliability and uptime of production systems.
  • Respond to incidents and conduct post-mortems to prevent recurrence.
  • Define and track SLOs/SLIs and improve capacity planning.
  • Collaborate with development teams to embed reliability in services.

Skills

Kubernetes
SRE principles
Incident response

Education

Bachelor's degree in Computer Science

Tools

Helm
Prometheus
Grafana
Terraform
OpenTelemetry
GitHub Actions

Job description

FactSet creates flexible, open data and software solutions for over 200,000 investment professionals worldwide, providing instant access to financial data and analytics that investors use to make crucial decisions.At FactSet, our values are the foundation of everything we do. They express how we act and operate, serve as a compass in our decision-making, and play a big role in how we treat each other, our clients, and our communities. We believe that the best ideas can come from anyone, anywhere, at any time, and that curiosity is the key to anticipating our clients’ needs and exceeding their expectations.**About the Role**We are looking for a skilled and motivated **Lead Site Reliability Engineer** to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our systems and services. You will work closely with development and operations teams to build and maintain robust infrastructure, automate processes, and drive engineering best practices. **Key Responsibilities*** Monitor, maintain, and improve the reliability and availability of production systems* Respond to and resolve incidents, conducting thorough post-mortems to prevent recurrence* Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)* Collaborate with development teams to build reliability into services from the ground up* Design and implement automation to reduce toil and improve operational efficiency* Participate in an on-call rotation to support critical systems* Contribute to capacity planning and performance optimization efforts* Document systems, processes, and runbooks to support the wider team **Required Technical Skills**:**Kubernetes** *(Required)** Hands-on experience deploying, managing, and troubleshooting workloads in Kubernetes* Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and Ingress* Experience with Kubernetes cluster management and administration* Familiarity with Helm for application packaging and deployment* Understanding of Kubernetes networking, storage, and security best practices* Bachelors degree in computer science or relevant degree.* Willing to work a hybrid model* Must be fluent in English both verbal and written* undefined **Additional Technical Skills*** **Cloud Platforms:** *(e.g. AWS, GCP, Azure)** **CI/CD Tooling:** *(e.g. GitHub Actions, ArgoCD, Harness)** **Monitoring & Observability:** *(e.g. Prometheus, Grafana, Coralogix, OpenTelemetry)** **Infrastructure as Code:** *(e.g. Terraform, Pulumi)** **Config Management**: *(e.g. Ansible, Puppet, Chef)** **Programming/Scripting:** *(e.g. Python, Go, Bash)* **Soft Skills & General Requirements*** Strong problem-solving and analytical skills with a methodical approach to troubleshooting* Excellent communication skills with the ability to collaborate across technical and non-technical teams* A proactive mindset with a focus on automation and continuous improvement* Ability to work effectively under pressure, particularly during incident response* Commitment to a blameless culture and continuous learning**Nice to Have*** Experience contributing to open-source projects* Familiarity with SRE principles as defined by the Google SRE handbook* Previous experience in a DevOps or Platform Engineering role **Company Overview:**FactSet (NYSE:FDS | NASDAQ:FDS) helps the financial community to see more, think bigger, and work better. Our digital platform and enterprise solutions deliver financial data, analytics, and open technology to more than 8,200 global clients, including over 200,000 individual users. Clients across the buy-side and sell-side, as well as wealth managers, private equity firms, and corporations, achieve more every day with our comprehensive and connected content, flexible next-generation workflow solutions, and client-centric specialized support. As a member of the S&P 500, we are committed to sustainable growth and have been recognized among the Best Places to Work in 2023 by Glassdoor as a Glassdoor Employees’ Choice Award winner. Learn more at www.factset.com and follow us on X and LinkedIn. At FactSet, we celebrate difference of thought, experience, and perspective. Qualified applicants will be considered for employment without regard to characteristics protected by law.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (Kubernetes Required) - Hybrid
Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

20035 FactSet Europe Limited • Greater London

Hybrid
GBP 90,000 - 150,000
Senior Site Reliability Engineer - Kubernetes & Automation
Senior Site Reliability Engineer - Kubernetes & Automation

20035 FactSet Europe Limited • Greater London

Hybrid
GBP 90,000 - 150,000
Sales Specialist - Real Time
Sales Specialist - Real Time

FactSet • Greater London

On-site
GBP 80,000 - 120,000
Health insurance
Life and disability insurance
Retirement savings plans
+1
Sales Specialist - Real Time
Sales Specialist - Real Time

20035 FactSet Europe Limited • Greater London

On-site
GBP 60,000 - 120,000
Health insurance
Retirement savings plan
Employee stock purchase program
+1
Assistant Manager, Media Production
Assistant Manager, Media Production

FactSet Research Systems Inc. • Greater London

Hybrid
GBP 45,000 - 65,000
Health insurance
Retirement savings plans
Employee stock purchase program
+2
Principal Sales Executive- Trading (Portware)
Principal Sales Executive- Trading (Portware)

FactSet Research Systems Inc. • Greater London

On-site
GBP 120,000 - 190,000
Health insurance
Life insurance
Disability insurance
+3
Senior Director of Agentic Software Development
Senior Director of Agentic Software Development

FactSet • Greater London

On-site
GBP 150,000 - 210,000
Business Development Representative
Business Development Representative

20035 FactSet Europe Limited • Greater London

On-site
GBP 35,000 - 55,000
Health insurance
Life insurance
Disability insurance
+3
Content Manager - Corporate Actions
Content Manager - Corporate Actions

FactSet • Greater London

Hybrid
GBP 72,000 - 89,000
Health, life & disability insurance
Retirement savings plans
Employee stock purchase program
+4
Lead SRE: Build Reliable, Scalable Systems
Lead SRE: Build Reliable, Scalable Systems

FactSet Research Systems Inc. • Greater London

Hybrid
GBP 90,000 - 130,000