Senior Site Reliability Engineer (Platform Reliability, Resilience)

Elastic

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health coverage for you and family
Flexible location & schedule
Generous vacation days
16+ weeks parental leave
Volunteer time
Charitable giving matched

Job summary

Elastic is seeking an experienced Site Reliability Engineer within the Platform Engineering team to design, build, and scale a multi-cloud platform that hosts Elastic Cloud Hosted and Serverless. You will automate system engineering tasks to guarantee global reliability and participate in follow-the-sun on-call rotations.

You will work with Golang, Kubernetes, and IaC tools to extend tooling, monitor performance, and improve incident response across distributed teams and environments.

Qualifications

  • Experience building and operating SaaS platforms in public cloud environments.
  • Strong programming skills in Golang or similar languages.
  • Hands-on experience with Kubernetes across multiple clouds.
  • Expertise in alerting and major incident management practices.
  • Familiarity with IaC tooling (Terraform, Crossplane) and containerized services.

Responsibilities

  • Lead automation to guarantee reliability of Elastic infrastructure.
  • Participate in on-call rotations with follow-the-sun coverage.
  • Respond to incidents, perform root cause analyses and implement durable fixes.
  • Design and scale multi-cloud platform components for Elastic Cloud Hosted and Serverless.
  • Collaborate across distributed teams to elevate reliability and performance.

Skills

Golang
Kubernetes
Terraform
Crossplane
Linux
Public cloud
Incident management
Alerting
On-call
Distributed teams

Tools

Elastic Stack
Prometheus
Graphite
Influx
Docker

Job description

  • As part of the Platform Engineering department, the SRE team is designing, building, scaling and maturing the multi-cloud platform for hosting internal and external services such as the Elastic Cloud Hosted and Serverless
  • We develop and extend new software and tools that support the rest of the infrastructure, so that we can rapidly deploy products from all corners of Elastic
  • We want your experience and recommendations to offer a truly exceptional customer experience!
  • Taking an engineering approach in leading technical initiatives for automating system engineering efforts to guarantee the reliability of the global Elastic infrastructure
  • Growing our global Platform infrastructure to meet the increasing scaling demands by developing and maintaining software, tooling and automations
  • Using an inclusive approach at championing an environment focused on collaboration, operational excellence, and uplifting others
  • Responding to and preventing repeated customer impact in response to major incidents and prioritised problem management. Our on call rotation uses follow-the-sun model where everyone participates in it in (mostly) their working hours
Benefits
  • Toast to your health: Fully paid health coverage for you and your family, in many locations.
  • Craft your calendar: Flexible location and schedule for most roles.
  • Create space for you: Distributed by design workforce, plus generous number of vacation days each year.
  • Embrace parenthood: Minimum of 16 weeks of parental leave, plus generous family formation benefits.
  • Give back your time: 40 hours each year to use toward volunteering with organizations and causes you’re passionate about.
  • Amplify your impact: Double your charitable giving — we match donations up to $1500 USD (or local currency equivalent).

Passion for developing solutions that involve inclusive communication methods to grow and strengthen partner and team relationships. Examples of working in distributed teams or working remotely is desirableA background in software engineering to collaborate with engineers to expertly identify, implement and deliver solutions. An experience in public cloud and managed Kubernetes services is advantageousSuccess and lessons of experiences from striving for ‘progress not perfection’ in the name of Platform reliability. We want to hear about your customer first approach in solving operational problems with a SRE perspectiveYou have operated a SaaS product in a public cloud ideally built using Infrastructure-as-Code tooling such as Crossplane or TerraformYou have built or operated a Kubernetes-at-scale infrastructure, ideally across multiple cloud providers, and the vital automation to support itYou have written non-trivial programs in Golang or other programming languagesYou have worked with containerized services (such as Docker.)You have proven experience in leading and improving alerting and major incident management standard processes metrics systems (e.g. Elastic Stack, Graphite, Prometheus, Influx) to diagnose issues and quantify impacts to present to others at varying level of the organizationYou have experience in system administration with professional skills in Linux on distributed systems at scaleYou have diagnosed or designed, implemented and created solutions with the Elastic StackYou are experienced in thriving in a self-organizing and sharing in a globally distributed team environmentYou strengthen team members in bringing out the best of each other by uplifting others with coaching and mentoring

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform SRE: Multi-Cloud Reliability & Resilience
Senior Platform SRE: Multi-Cloud Reliability & Resilience

Elastic • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage for you and family
Flexible location & schedule
Generous vacation days
+3
Senior SRE
Senior SRE

Pulse Recruit • Greater London

On-site
GBP 65,000 - 85,000
SRE Engineer
SRE Engineer

Selby Jennings • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

On-site
GBP 51,000 - 85,000
Bonus
Benefits
Backend Engineer
Backend Engineer

TESTQ Technologies LTD. • Sheffield

Hybrid
GBP 70,000 - 110,000
Senior Cloud SRE: Multi-Cloud Reliability & Automation
Senior Cloud SRE: Multi-Cloud Reliability & Automation

Elasticsearch B.V. • United Kingdom

Remote
GBP 75,000 - 140,000
Health coverage
Flexible locations and schedules
Generous vacation days
+4
Manager
Manager

Cameroon Mathematical Union (CAMU) • Greater London

Hybrid
GBP 15,000 - 38,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

iXceed Solutions • Basildon

On-site
GBP 55,000 - 75,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Factset • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage
Free lunch in the office (Mon–Fri)
Employee social events and sports