AI Observability Engineer: Build AI Health & Metrics

AI Chopping Block

Amsterdam

On-site

EUR 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Career growth and learning
Flexibility and ownership
Collaborative culture
Impactful AI projects
International environment

Job summary

Nebius is building a high-performance AI cloud platform in Amsterdam and seeks an AI Observability Engineer to own monitoring across LLMs, agents, and platform workloads. You will implement traces, latency, token usage, and cost analytics, while crafting Grafana dashboards and Prometheus metrics to ensure reliability and self-service visibility.

The role involves Terraform IaC, CI/CD for observability tooling, Python-based exporters, and collaboration with Platform Engineering to drive health,

Qualifications

  • 5-8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering.
  • Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana.
  • Hands-on experience with Langfuse, Grafana, and Prometheus.
  • Experience with Terraform and CI/CD.
  • Python skills for instrumentation, exporters, and automation.
  • Familiarity with ML workloads and AI-specific metrics.
  • Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs.
  • Intermediate or higher English.

Responsibilities

  • Stand up and operate LLM and agent monitoring with Langfuse.
  • Capture traces, latency, token usage, cost, quality scores, prompt and model-version analytics, and safety signals.
  • Build lightweight internal tooling and exporters in Python.
  • Design and maintain Grafana dashboards, Prometheus metrics, and the Azure observability stack.
  • Instrument platform and AI workloads for health, usage, cost, and SLA reporting.
  • Feed telemetry and operational insights into the Platform Engineering backlog.
  • Own Terraform IaC and CI/CD for observability tooling.
  • Support incident investigation and root-cause analysis.

Skills

Observability
SRE
Cloud engineering
CI/CD
Python
English
Azure Monitor
Application Insights
Log Analytics
Grafana
Prometheus

Tools

Langfuse
Grafana
Prometheus
Azure
Terraform
CI/CD tooling

Job description

Nebius is building a high-performance AI cloud platform in Amsterdam and seeks an AI Observability Engineer to own monitoring across LLMs, agents, and platform workloads. You will implement traces, latency, token usage, and cost analytics, while crafting Grafana dashboards and Prometheus metrics to ensure reliability and self-service visibility.

The role involves Terraform IaC, CI/CD for observability tooling, Python-based exporters, and collaboration with Platform Engineering to drive health,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Observability Engineer — Build AI Platform Telemetry
AI Observability Engineer — Build AI Platform Telemetry

Nebius • Amsterdam

On-site
EUR 90,000 - 125,000
Competitive compensation
Career growth
Flexibility and ownership
+3
AI Observability Engineer: Azure Metrics & Tracing Lead
AI Observability Engineer: Azure Metrics & Tracing Lead

Nebius B.V. • Amsterdam

On-site
EUR 70,000 - 110,000
AI Observability Engineer
AI Observability Engineer

Nebius B.V. • Amsterdam

On-site
EUR 70,000 - 110,000
AI Observability Engineer
AI Observability Engineer

AI Chopping Block • Amsterdam

On-site
EUR 90,000 - 130,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
AI Observability Engineer
AI Observability Engineer

Nebius • Amsterdam

On-site
EUR 90,000 - 125,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Product Growth Analytics Lead — AI Cloud Platform
Product Growth Analytics Lead — AI Cloud Platform

Nebius • Amsterdam

On-site
EUR 90,000 - 130,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Remote AI/ML Solutions Architect for Cloud Deployments
Remote AI/ML Solutions Architect for Cloud Deployments

Meyandy LLC • Amsterdam

Hybrid
EUR 120,000 - 190,000
Competitive compensation
Career growth opportunities
Flexible work arrangement
+1
Head of Platform Engineering — Scale AI Infrastructure
Head of Platform Engineering — Scale AI Infrastructure

AI Chopping Block • Amsterdam

Hybrid
EUR 140,000 - 190,000
Competitive compensation
Career growth and learning
Ownership and autonomy
+1
Senior Full-Stack AI Engineer - Production AI Apps
Senior Full-Stack AI Engineer - Production AI Apps

Nebul • Leiden

On-site
EUR 90,000 - 130,000
Office near The Hague
Senior Full-Stack AI Engineer (LLMs & Agentic Systems)
Senior Full-Stack AI Engineer (LLMs & Agentic Systems)

Nebul • Leiden

On-site
EUR 90,000 - 130,000
Office near The Hague