Enable job alerts via email!

Senior Site Reliability Engineer, Production Engineering

ThousandEyes (part of Cisco)

London

Hybrid

GBP 70,000 - 90,000

Full time

Yesterday
Be an early applicant

Boost your interview chances

Create a job specific, tailored resume for higher success rate.

Job summary

A leading company in the Digital Experience Assurance space is seeking a Senior Site Reliability Engineer in London. The role involves designing and managing large-scale distributed systems, collaborating with development teams to enhance platform reliability and performance. Candidates should possess expert knowledge of Kubernetes, proficiency in programming languages like Python or Go, and a strong understanding of cloud services, particularly AWS. This hybrid position requires at least one day a week in the London office.

Qualifications

  • Expert-level knowledge of Kubernetes and its ecosystem.
  • Proficiency in software development with languages such as Python or Go.
  • 5+ years of experience in a related role.

Responsibilities

  • Collaborate with software engineers to optimize architecture and services.
  • Design, deploy, and maintain AWS cloud-native services.
  • Automate production operations for continuous platform operation.

Skills

Kubernetes
Python
Go
Unix/Linux
Incident Response
Change Management
Distributed Systems

Job description

Senior Site Reliability Engineer, Production Engineering
Please note that we have a hybrid approach to work and would like to find someone who can come into our offices in London at least one day a week.
Who We Are

Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences.

ThousandEyes is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios.

About The Role

We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.

What You’ll Do
  • Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
  • Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
  • Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
  • Participate in and improve our 24x7 incident response and on-call rotation.
  • Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
  • Automate production operations to provide guardrails and continuous platform operation.
  • Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
  • Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
  • Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
  • Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
  • Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
  • Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.
Qualifications
  • Expert-level knowledge of Kubernetes and its ecosystem.
  • Proficiency in software development with languages such as Python or Go.
  • In-depth knowledge of cloud providers, preferably AWS.
  • Proven ability to build and implement scalable and well-tested solutions.
  • Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols.
  • Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs.
Preferred Qualifications
  • Familiarity with best practices for operating a large-scale, highly available enterprise platform.
  • 5+ years of experience in a related role.
  • Excellent communication and documentation skills.
  • Strong sense of ownership, drive, and attention to detail.

Cisco values the perspectives and skills that emerge from employees with diverse backgrounds. That's why Cisco is expanding the boundaries of discovering top talent by not only focusing on candidates with educational degrees and experience but also placing more emphasis on unlocking potential. We believe that everyone has something to offer and that diverse teams are better equipped to solve problems, innovate, and create a positive impact.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification. Research shows that people from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy. We urge you not to prematurely exclude yourself and to apply if you're interested in this work.

Apply for this job

*

indicates a required field

First Name *

Last Name *

Preferred First Name

Email *

Phone *

Location (City) *

Resume/CV *

Enter manually

Accepted file types: pdf, doc, docx, txt, rtf

Enter manually

Accepted file types: pdf, doc, docx, txt, rtf

Education

School Select...

Degree Select...

Select...

LinkedIn Profile

Website

How did you hear about this job?

Are you now legally authorized to work in the posted primary location for this requisition? * Select...

Will you require sponsorship in the future for this location (for example, if you are on a temporary visa)? * Select...

Get your free, confidential resume review.
or drag and drop a PDF, DOC, DOCX, ODT, or PAGES file up to 5MB.

Similar jobs

Senior Site Reliability Engineer, Production Engineering New London, Greater London, England, U[...]

ThousandEyes

London

Hybrid

GBP 70,000 - 100,000

Yesterday
Be an early applicant

Senior Site Reliability Engineer, Production Engineering

ThousandEyes

London

Hybrid

GBP 70,000 - 90,000

Today
Be an early applicant

Site Reliability Engineer (Home-based)

JR United Kingdom

London

Remote

GBP 60,000 - 80,000

4 days ago
Be an early applicant

Senior Site Reliability Engineer

Auros

Greater London

Remote

GBP 60,000 - 100,000

16 days ago

Senior Site Reliability Engineer

Prima

London

Hybrid

GBP 70,000 - 90,000

Today
Be an early applicant

Senior Site Reliability Engineer

Thomson Reuters

London

On-site

GBP 70,000 - 90,000

Today
Be an early applicant

Senior Site Reliability Engineer - AWS Kubernetes

Source Technology

London

On-site

GBP 70,000 - 90,000

Today
Be an early applicant

Senior Site Reliability Engineer

TRSS

London

On-site

GBP 70,000 - 90,000

Today
Be an early applicant

Remote Senior Site Reliability Engineer Manager (Remote)

Remotestar

London

Remote

GBP 80,000 - 100,000

30+ days ago