Senior Site Reliability Engineer

Jobgether

United Kingdom

On-site

GBP 64,000 - 85,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote
Global cross‑functional team
Ownership of infrastructure

Job summary

Jobgether is seeking a Senior Site Reliability Engineer for a fully remote role focused on cloud infrastructure, reliability, and observability for AI/ML geospatial projects. You will own GCP infrastructure, drive SLOs, incident response, and performance improvements across distributed teams.

Based in Canada, you will collaborate with Product and Engineering across North America, Europe, and the UK, applying Terraform, Kubernetes, and OpenTelemetry to deliver scalable cloud systems with cost

Qualifications

  • 3+ years of professional experience in SRE/production engineering or related infra role.
  • Hands-on experience with Google Cloud Platform, cost governance and infra management.
  • Proven Kubernetes experience in production environments.
  • Experience with Infrastructure as Code tools such as Terraform or Deployment Manager.
  • Scripting skills in Python, Bash or Go.
  • Observability expertise with Prometheus, Grafana, OpenTelemetry and logging systems.
  • Experience with incident management, post-incident reviews, production troubleshooting, and on-call operations.
  • Ability to define, monitor and improve SLOs, SLIs and error budgets.

Responsibilities

  • Design, build, and evolve scalable cloud infrastructure on Google Cloud Platform.
  • Develop internal tooling, automation, and self-service capabilities to boost engineering efficiency.
  • Strengthen the observability platform across metrics, logging, tracing, and monitoring.
  • Establish cost visibility and governance to optimize cloud spending.
  • Define and champion reliability practices including SLOs, SLIs, error budgets, and DORA metrics.
  • Lead incident management, coordinate responses, and drive improvements from post-incident reviews.
  • Participate in on-call rotation to keep production systems stable and reliable.
  • Collaborate with Product and Engineering to improve deployment reliability and software delivery.

Skills

Google Cloud Platform
Kubernetes
Scripting: Python/Bash/Go
Observability tools
SLO/SLI and error budgets
Incident management
DORA metrics
On-call readiness

Tools

Terraform
Deployment Manager
OpenTelemetry

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Canada.

Join a growing AI/ML organization operating at the intersection of geospatial intelligence and environmental technology. As a Senior Site Reliability Engineer, you will take ownership of cloud infrastructure and help strengthen reliability, observability, and operational excellence across engineering and product teams. You will design and evolve scalable infrastructure on Google Cloud Platform while enabling developers through automation and self-service tooling. Your work will directly influence system availability, incident response, deployment performance, and cloud efficiency. You will champion modern reliability practices, from SLOs and error budgets to DORA metrics and observability. This is a fully remote opportunity offering significant technical ownership in a collaborative, high-impact engineering environment.

Accountabilities
  • Design, build, and continuously evolve scalable and reliable cloud infrastructure on Google Cloud Platform.
  • Develop internal tooling, automation, and self-service capabilities that increase engineering efficiency and reduce operational dependencies.
  • Strengthen the observability platform across metrics, logging, distributed tracing, and monitoring to improve system visibility and reduce mean time to recovery.
  • Establish infrastructure cost visibility, governance, and optimization initiatives to improve cloud efficiency and manage spending responsibly.
  • Define and champion reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and DORA metrics.
  • Lead incident management activities, coordinate response efforts, facilitate post-incident reviews, and drive meaningful improvements based on incident learnings.
  • Participate in the on-call rotation and help ensure production systems remain stable, available, and resilient.
  • Partner with Product and Engineering teams to improve operational practices, deployment reliability, and overall software delivery performance.
Requirements
  • 3+ years of professional experience in Site Reliability Engineering, production engineering, or a closely related infrastructure role.
  • Strong hands-on experience with Google Cloud Platform, including cloud cost optimization, governance, and infrastructure management.
  • Proven experience managing Kubernetes clusters and workloads in production environments.
  • Experience with Infrastructure as Code tools such as Terraform or Google Cloud Deployment Manager.
  • Strong scripting and automation capabilities using Python, Bash, Go, or comparable languages.
  • Extensive experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, centralized logging, and distributed tracing.
  • Practical experience with incident management, post-incident analysis, production troubleshooting, and on-call operations.
  • Ability to define, implement, and monitor SLOs, SLIs, and error budgets.
  • Strong understanding of reliability engineering principles and a proactive approach to identifying and resolving operational risks.
  • Familiarity with DORA metrics and their application to engineering workflows is an asset.
  • Experience working in AI/ML, geospatial technology, or other data-intensive technical environments is considered a plus.
  • Strong communication and collaboration skills, with the ability to work effectively across distributed Product and Engineering teams.
  • Candidates must be authorized to work in their country of residence; visa sponsorship is not available.
Benefits
  • Fully remote work environment.
  • Opportunity to work with modern cloud infrastructure, Kubernetes, Infrastructure as Code, and advanced observability technologies.
  • Significant ownership over infrastructure reliability, operational excellence, and cloud optimization initiatives.
  • Opportunity to influence engineering practices through SLOs, error budgets, DORA metrics, and incident management.
  • Collaboration with cross-functional and distributed engineering teams across North America, Europe, and the UK.
  • Exposure to innovative AI/ML and geospatial technology applications.
  • Compensation will be discussed during the interview process and will be aligned with experience and geographic location.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE, Remote Cloud Reliability for AI/Geospatial
Senior SRE, Remote Cloud Reliability for AI/Geospatial

Jobgether • United Kingdom

On-site
GBP 64,000 - 85,000
Fully remote
Global cross‑functional team
Ownership of infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

P2P • Greater London

On-site
GBP 90,000 - 130,000
Site Reliability Engineer (SRE) / Platform Engineer
Site Reliability Engineer (SRE) / Platform Engineer

Adecco • City Of London

Hybrid
Hybrid work arrangement
Competitive day rate
London-based contract
Site Reliability Engineer
Site Reliability Engineer

WALT Labs • Greater London

On-site
GBP 70,000 - 90,000
20 holiday days + bank holidays
Mentoring and training opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Senior Infrastructure Operations Engineer, Sovereign Operations, Wiltshire
Senior Infrastructure Operations Engineer, Sovereign Operations, Wiltshire

Google Inc. • Greater London

On-site
GBP 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Understanding Recruitment • Greater London

On-site
GBP 150,000 - 200,000
UK visa sponsorship
Equity
Private healthcare
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Symphony • Belfast City District

On-site
GBP 60,000 - 70,000
Regional specific competitive benefits
Build your own Benefits (BYOB) perk
Local events, team building, and devop
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

Hybrid
GBP 140,000 - 200,000
ESPP
Life assurance
Income protection
+11