Observability Platform SRE - Greenfield, Cloud-Native

Aalyria

Greater London

Hybrid

GBP 110,000 - 150,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity participation
Pension plan
Private health insurance
Hybrid work model

Job summary

Aalyria is seeking an experienced Site Reliability/Platform Engineer to build the core observability stack for satellite and deep-space platforms. You will design centralized monitoring across metrics, logs, and traces and drive reliability through SLOs/SLIs and incident response.

You will collaborate with SWE teams to implement best practices, templates, and tooling, while maturing the observability stack with Terraform and ArgoCD in GCP/AWS environments.

Qualifications

  • 4+ years in an SRE or platform engineering role with a focus on observability for large-scale, distributed compute or network systems.
  • Hands-on experience building, scaling, and managing observability platforms (Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems.
  • Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes.
  • Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD).
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and tooling.
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services.

Responsibilities

  • Design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (Prometheus), logging (Loki), and distributed tracing (Tempo/OpenTelemetry).
  • Define, implement, and manage a robust framework of SLOs, SLIs, and error budgets for core products.
  • Partner with SWEs to implement observability best practices, templates, documentation, and tooling configuration.
  • Automate deployment, scaling, and management of the observability stack using Terraform and GitOps (ArgoCD).
  • Collaborate with core infrastructure to ensure visibility into Kubernetes clusters and cloud environments (GCP and AWS).
  • Develop and lead monitoring, alerting, and incident response strategies with blameless post-mortems.

Skills

Prometheus
Grafana
Loki/ELK
OpenTelemetry
Tempo/Jaeger
Kubernetes
Go/Python tooling

Tools

Terraform
ArgoCD
GitLab CI
Istio/Linkerd
Go
Python

Job description

Aalyria is seeking an experienced Site Reliability/Platform Engineer to build the core observability stack for satellite and deep-space platforms. You will design centralized monitoring across metrics, logs, and traces and drive reliability through SLOs/SLIs and incident response.

You will collaborate with SWE teams to implement best practices, templates, and tooling, while maturing the observability stack with Terraform and ArgoCD in GCP/AWS environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform SRE for Observability — Satellite Networks
Platform SRE for Observability — Satellite Networks

Aalyria Technologies, Inc. • Greater London

Hybrid
GBP 90,000 - 130,000
Equity participation
Private health insurance
Generous leave
Site Reliability Engineer- Spacetime UK
Site Reliability Engineer- Spacetime UK

Aalyria Technologies, Inc. • Greater London

Hybrid
GBP 90,000 - 130,000
Equity participation
Private health insurance
Generous leave
Site Reliability Engineer- Spacetime UK
Site Reliability Engineer- Spacetime UK

Aalyria • Greater London

Hybrid
GBP 110,000 - 150,000
Equity participation
Pension plan
Private health insurance
+1
Observability SRE AVP: Cloud Observability & Migrations
Observability SRE AVP: Cloud Observability & Migrations

Citi • Greater London

On-site
GBP 90,000 - 120,000
Senior SRE, Observability & Cloud Reliability
Senior SRE, Observability & Cloud Reliability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Omilia • United Kingdom

On-site
GBP 90,000 - 130,000
Fixed compensation
Long-term employment
Professional growth
+1
Observability AVP: SRE & Cloud Telemetry Lead
Observability AVP: SRE & Cloud Telemetry Lead

Citigroup Inc. • Greater London

On-site
GBP 120,000 - 180,000
Senior Site Reliability Engineer - Cloud Observability & Automation
Senior Site Reliability Engineer - Cloud Observability & Automation

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
Observability SRE — Reliability & Telemetry Engineer, London
Observability SRE — Reliability & Telemetry Engineer, London

HCLTech • Greater London

On-site
GBP 70,000 - 95,000
AI Platform SRE - Infra, Security & Observability
AI Platform SRE - Infra, Security & Observability

duvo.ai • United Kingdom

Remote
GBP 70,000 - 120,000
Unlimited AI budget
Autonomy to do your best work
Real AI product with real customers
+2