Senior Platform & Observability Engineer

Koda Tech

Greater London

Hybrid

GBP 90,000 - 120,000

Full time

11 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Koda Tech in London is seeking a Senior Platform & Observability Engineer to shape and mature our observability strategy for large-scale, mission‑critical systems.

You will define standards across engineering, infrastructure and support, improve incident response, and design a target-state architecture with telemetry, logging and tracing. The role is hybrid, combining hands-on delivery with technical leadership.

Qualifications

  • Strong experience solving real operational problems rather than focusing on tools.
  • Experience defining monitoring, logging or reliability standards.
  • Experience working across engineering, infrastructure, and support teams.
  • Deep understanding of metrics, logs, alerting and service health concepts.
  • Experience scaling structured logging practices.
  • Strong Infrastructure as Code experience.
  • Experience in hybrid cloud and on-prem environments.

Responsibilities

  • Assess the current observability landscape across applications, infrastructure and services.
  • Identify gaps in monitoring, alerting and operational workflows.
  • Define observability standards and best practices across the business.
  • Design a target-state observability architecture and roadmap.
  • Work with engineers to improve telemetry, metrics, logging and tracing.
  • Drive adoption of structured logging practices.
  • Establish standards for alerting, incident detection and service health monitoring.
  • Collaborate with platform, infrastructure and operations teams to automate observability via IaC.
  • Act as a trusted advisor on reliability, monitoring and operational excellence.
  • Support the wider Platform and DevOps function beyond the observability programme.

Skills

Site Reliability Engineering
Observability Engineering
Cloud architecture

Tools

OpenTelemetry
Grafana
Elastic
Splunk
Terraform

Job description

London | Hybrid (3 days per week in office)

We're hiring a Senior Platform & Observability Engineer to help shape and mature the observability strategy for a growing technology organisation operating large-scale business-critical systems.

This is an opportunity for someone who enjoys solving complex operational challenges, influencing technical direction, and building practical solutions that improve reliability, visibility and engineering effectiveness.

Rather than simply maintaining monitoring tools, you'll help define how observability should work across the organisation, working closely with engineering, infrastructure and support teams to establish standards, improve incident response, and create a clearer picture of system health.

What you'll be doing
  • Assessing the current monitoring and observability landscape across applications, infrastructure and services
  • Identifying gaps in existing monitoring, alerting and operational workflows
  • Defining observability standards and best practices across the business
  • Designing a target-state observability architecture and roadmap
  • Working with software engineers to improve telemetry, metrics, logging and tracing
  • Driving adoption of structured logging practices
  • Helping establish standards for alerting, incident detection and service health monitoring
  • Collaborating with platform, infrastructure and operations teams to automate observability capabilities through Infrastructure as Code
  • Acting as a trusted advisor on reliability, monitoring and operational excellence
  • Supporting the wider Platform and DevOps function beyond the observability programme
What we're looking for

We're more interested in experience solving real operational problems than specific tools.

You'll likely have experience in several of the following:

  • Site Reliability Engineering (SRE)
  • Observability Engineering
We'd love to speak with people who have:
  • Built or significantly improved observability capabilities within a business
  • Defined monitoring, logging or reliability standards rather than simply operating existing tools
  • Experience working across engineering, infrastructure and support teams
  • Strong understanding of metrics, logs, alerting and service health concepts
  • Experience introducing or scaling structured logging practices
  • Strong Infrastructure as Code experience
  • Experience working with hybrid environments, including both cloud and physical infrastructure, is highly desirable
  • Confidence challenging existing approaches and driving technical change
Nice to have

Experience with technologies such as:

  • OpenTelemetry
  • Grafana
  • Elastic
  • Splunk
  • Terraform
  • Public cloud platforms

No single technology is required. We're interested in how you've used observability to solve problems rather than which tools you've worked with.

The type of person who succeeds here

This role would suit someone who has worked in a startup or scale-up environment and has helped build processes, standards or platforms from the ground up.

You'll be comfortable navigating ambiguity, influencing technical decisions, and helping teams move towards a more mature operational model.

If you've previously inherited a fragmented monitoring landscape and successfully implemented a clearer, more effective approach to observability, we'd like to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

Hybrid
GBP 51,000 - 85,000
Bonus
Benefits
Observability SRE
Observability SRE

HCLTech • Greater London

On-site
GBP 70,000 - 95,000
Observability Engineer
Observability Engineer

La-Fosse-1 • Liverpool

On-site
GBP 60,000 - 90,000
Competitive salary
Annual leave
Pension
+4
Site Reliability Engineer
Site Reliability Engineer

La Fosse • Liverpool

On-site
GBP 60,000 - 90,000
Competitive package
Generous annual leave
Pension
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Senior Platform Engineer
Senior Platform Engineer

KDR Talent Solutions • Greater London

Hybrid
GBP 120,000 - 150,000
Platform Engineer
Platform Engineer

Albert Bow • Greater London

On-site
GBP 120,000 - 150,000
Observability Engineer
Observability Engineer

La Fosse Associates • Liverpool

On-site
GBP 65,000 - 90,000
Competitive salary
Generous annual leave
Excellent pension contribution
+4
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Platform Engineer (DevOps & Kubernetes)
Platform Engineer (DevOps & Kubernetes)

Be-IT • Greater London

Hybrid
GBP 80,000 - 110,000
Hybrid working in Central London