AI Observability Engineer

Nebius B.V.

Amsterdam

On-site

EUR 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius is building a high-performance AI cloud platform and seeks an AI Observability Engineer to own the observability backbone—from LLM and agent tracing to platform metrics—so the Azure AI platform is visible, measurable, reliable, and self-service.

You will stand up monitoring with Langfuse, capture traces, latency, token usage, cost, quality scores, and AI metrics, and design Grafana dashboards while owning Terraform IaC and CI/CD for observability tooling.

Qualifications

  • 5–8 years in observability, SRE, or cloud/platform engineering.
  • Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana.
  • Hands-on with Langfuse, Grafana, and Prometheus.
  • Experience with Terraform and CI/CD.
  • Python skills for instrumentation and automation.
  • Familiarity with ML workloads and AI metrics.
  • English at intermediate or higher.

Responsibilities

  • Stand up and operate LLM and agent monitoring with Langfuse.
  • Capture traces, latency, token usage, cost, quality scores, and AI metrics.
  • Build lightweight internal tooling and exporters in Python.
  • Design and maintain Grafana dashboards, Prometheus metrics, and Azure observability stack.
  • Instrument platform and AI workloads for health, usage, cost, and SLA reporting.
  • Feed telemetry and operational insights into the Platform Engineering backlog.
  • Own Terraform IaC and CI/CD for observability tooling.
  • Support incident investigation and root-cause analysis.

Skills

English (intermediate)

Tools

Azure Monitor
Application Insights
Log Analytics
Managed Grafana
Langfuse
Grafana
Prometheus
Terraform
CI/CD
Python
Kusto Query Language
PromQL

Job description

Nebius is building a high-performance AI cloud platform, and we are looking for an AI Observability Engineer to own the AI observability backbone—from LLM and agent tracing to platform and infrastructure metrics—so the Azure AI platform is visible, measurable, reliable, and self-service.

Your responsibilities:
  • Stand up and operate LLM and agent monitoring with Langfuse.
  • Capture traces, latency, token usage, cost, quality scores, prompt and model-version analytics, and safety signals.
  • Build lightweight internal tooling and exporters in Python.
  • Design and maintain Grafana dashboards, Prometheus metrics, and the Azure observability stack.
  • Instrument platform and AI workloads for health, usage, cost, and SLA reporting.
  • Feed telemetry and operational insights into the Platform Engineering backlog.
  • Own Terraform IaC and CI/CD for observability tooling.
  • Support incident investigation and root-cause analysis.
Must-haves:
  • 5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering.
  • Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana.
  • Hands-on experience with Langfuse, Grafana, and Prometheus.
  • Experience with Terraform and CI/CD.
  • Python skills for instrumentation, exporters, and automation.
  • Familiarity with ML workloads and AI-specific metrics.
  • Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs.
  • Intermediate or higher English.
Nice-to-haves:
  • PromQL and Kusto Query Language.
  • OpenTelemetry, including GenAI semantic conventions.
  • LLM evaluation frameworks.
  • AI cost dashboards and FinOps.
  • Alerting, on-call, and incident management tooling.
  • AKS and Kubernetes observability.

5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering, Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana, Hands-on experience with Langfuse, Grafana, and Prometheus, Experience with Terraform and CI/CD, Python skills for instrumentation, exporters, and automation, Familiarity with ML workloads and AI-specific metrics, Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs, Intermediate or higher English

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Observability Engineer — Build AI Platform Telemetry
AI Observability Engineer — Build AI Platform Telemetry

Nebius • Amsterdam

On-site
EUR 90,000 - 125,000
Competitive compensation
Career growth
Flexibility and ownership
+3
AI Observability Engineer: Azure Metrics & Tracing Lead
AI Observability Engineer: Azure Metrics & Tracing Lead

Nebius B.V. • Amsterdam

On-site
EUR 70,000 - 110,000
AI Observability Engineer
AI Observability Engineer

AI Chopping Block • Amsterdam

On-site
EUR 90,000 - 130,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
AI Observability Engineer
AI Observability Engineer

Nebius • Amsterdam

On-site
EUR 90,000 - 125,000
Competitive compensation
Career growth
Flexibility and ownership
+3
AI Observability Engineer: Build AI Health & Metrics
AI Observability Engineer: Build AI Health & Metrics

AI Chopping Block • Amsterdam

On-site
EUR 90,000 - 130,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
AI Platform Engineer
AI Platform Engineer

Kyndryl • Hoofddorp

On-site
EUR 90,000 - 120,000
AI Governance & Platform Engineer (ID: 3863)
AI Governance & Platform Engineer (ID: 3863)

Stafide • Netherlands

Hybrid
EUR 70,000 - 90,000
Growth opportunities
Exposure to enterprise AI projects
Collaboration with tech teams
AI Architect
AI Architect

VBeyond Corporation • Den Haag

On-site
EUR 90,000 - 120,000
Senior AI Platform Engineer / AI Architect - Microsoft Technologies
Senior AI Platform Engineer / AI Architect - Microsoft Technologies

EPAM Systems • Netherlands

On-site
EUR 70,000 - 90,000
Senior AI Architect - Microsoft Technologies
Senior AI Architect - Microsoft Technologies

EPAM Systems • Netherlands

On-site
EUR 80,000 - 110,000