Stand out for this role — generate a tailored resume and cover letter in about a minute.
Factset is seeking a Lead Site Reliability Engineer to ensure reliability, scalability, and performance of critical systems. You will collaborate with development and operations to automate processes and embed reliability into services from design through deployment.
The role involves defining SLOs/SLIs, leading incident response, and driving continuous improvements across teams in a hybrid working model in the Greater London area.
Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and IngressHands-on experience deploying, managing, and troubleshooting workloads in KubernetesMust be fluent in English both verbal and writtenUnderstanding of Kubernetes networking, storage, and security best practicesFamiliarity with Helm for application packaging and deploymentBachelors degree in computer science or relevant degreeExperience with Kubernetes cluster management and administrationWilling to work a hybrid modelCommitment to a blameless culture and continuous learningStrong problem-solving and analytical skills with a methodical approach to troubleshootingExcellent communication skills with the ability to collaborate across technical and non-technical teamsAbility to work effectively under pressure, particularly during incident responseA proactive mindset with a focus on automation and continuous improvementExperience contributing to open-source projectsFamiliarity with SRE principles as defined by the Google SRE handbookPrevious experience in a DevOps or Platform Engineering role