Senior Software Engineer (LLM Ops & Evals)

Datasnipper

Netherlands

On-site

EUR 90,000 - 130,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Datasnipper in the Netherlands seeks an experienced backend/platform engineer to own the LLM gateway, multi-cloud deployments, and infra-as-code workflows. You will build and maintain a shared AI platform, ensuring reliability, security, and scalability while collaborating with product and ML teams.

You will lead on-call incidents, implement observability with OpenTelemetry and Grafana, and drive cost attribution across providers.

Qualifications

  • 5+ years of backend or platform engineering experience.
  • Strong production Python experience.
  • Experience with cloud infrastructure and IaC (Terraform).
  • Experience running shared services in production: on-call, incidents, postmortems, SLOs.
  • Experience building or running LLM inference infrastructure with a gateway or routing layer.

Responsibilities

  • Own the LLM gateway: routing, provider failover, rate limits, retries, and cost attribution across multiple model providers.
  • Deploy, version and deprecate models across clouds, regions and environments, including quota and capacity planning, managed as infrastructure as code.
  • Build the shared evaluation platform: versioned datasets, experiment tracking, run and result schemas, reporting, and trace linkage back to the run.
  • Own the infrastructure for async and long-running AI workloads.
  • Own observability for AI traffic: latency, retries and fallbacks, token usage, cost and errors, per team and per use case.
  • Take part in the on-call rotation, run incidents, and close the follow-ups.
  • Implement the security and compliance controls the platform is held to: retention, access control, RBAC and SSO.
  • Define and maintain clean integration contracts between the platform and the teams that consume it.
  • Partner with product and ML engineers to turn their requirements into platform capabilities that are self-service rather than a request queue.

Skills

Python
Observability
On-call
Collaboration
Judgment
Growth mindset

Tools

Terraform
Azure
GCP
OpenTelemetry
Grafana

Job description


  • Own the LLM gateway: routing, provider failover, rate limits, retries, and cost attribution across multiple model providers

  • Deploy, version and deprecate models across clouds, regions and environments, including quota and capacity planning, managed as infrastructure as code

  • Build the shared evaluation platform: versioned datasets, experiment tracking, run and result schemas, reporting, and trace linkage back to the run

  • Own the infrastructure for async and long-running AI workloads

  • Own observability for AI traffic: latency, retries and fallbacks, token usage, cost and errors, per team and per use case

  • Take part in the on-call rotation, run incidents, and close the follow-ups

  • Implement the security and compliance controls the platform is held to: retention, access control, RBAC and SSO

  • Define and maintain clean integration contracts between the platform and the teams that consume it

  • Partner with product and ML engineers to turn their requirements into platform capabilities that are self-service rather than a request queue



  • Comfort with privacy and compliance work: PII handling, anonymisation, retention, access control

  • Experience with cloud at the infrastructure level and infrastructure as code (we run across Azure and GCP with Terraform)

  • Experience building platform or shared-service capabilities consumed by multiple internal teams

  • 5+ years in backend or platform engineering, with strong production Python

  • Experience running a shared service in production: on-call, incidents, postmortems, SLOs

  • Experience building or running LLM inference infrastructure: a gateway or routing layer with multiple providers, failover, rate limiting and cost attribution

  • Hands-on experience with observability tools (OpenTelemetry, Grafana), including instrumenting services and designing dashboards and alerts

  • Judgment: You exercise sound judgment in ambiguous situations, balance speed and accuracy, and adjust priorities proactively

  • Ownership: You own work end-to-end, anticipate issues, and ensure high-quality delivery without close supervision

  • Adaptability: You navigate ambiguity calmly, model positive behavior, and help peers adjust through clear communication

  • Collaboration: You build strong cross-functional relationships and influence peers through expertise, data, and empathy

  • Growth Mindset: You encourage open feedback exchange and provide clear, balanced feedback that helps others grow

  • Synthetic data or document anonymisation pipelines

  • Document AI: VLMs, OCR, structured extraction and the metrics that go with it

  • Domain experience in audit, accounting or fintech

  • Temporal or another durable workflow engine

  • Self-hosted inference, capacity planning, load testing

  • Experience with LLM or agent evaluation

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

WCC Group • Utrecht

On-site
EUR 90,000 - 130,000
Hybrid work setup
International travel as needed
Machine Learning Engineer — LLMs, Retrieval and MLOps for NATO with security clearance
Machine Learning Engineer — LLMs, Retrieval and MLOps for NATO with security clearance

WLG • Den Haag

On-site
EUR 90,000 - 130,000
MLOps Engineer
MLOps Engineer

The French Sourcer • Amsterdam

On-site
EUR 90,000 - 130,000
Health insurance
Pension plan
Equity plan
+1
MLOps Engineer
MLOps Engineer

EPAM Systems • Netherlands

Hybrid
EUR 85,000 - 120,000
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Jobgether SRL • Netherlands

On-site
EUR 120,000 - 180,000
Competitive compensation
Career growth
Ownership over technical work
+1
Senior Applied AI Engineer (Agent Runtime (Alwin by DataSnipper)
Senior Applied AI Engineer (Agent Runtime (Alwin by DataSnipper)

Datasnipper • Netherlands

On-site
EUR 90,000 - 140,000
Forward Deployed Engineer, Benelux
Forward Deployed Engineer, Benelux

Telnyx • Amsterdam

On-site
EUR 120,000 - 160,000
Senior Software Engineer – Financial Statement Suite
Senior Software Engineer – Financial Statement Suite

Jobtailor • Amsterdam

On-site
EUR 90,000 - 130,000
Machine Learning Engineer — LLMs, Retrieval and MLOps for NATO with security clearance
Machine Learning Engineer — LLMs, Retrieval and MLOps for NATO with security clearance

Wlgroup • Den Haag

On-site
EUR 70,000 - 110,000
Senior AI / ML Engineer (MLOps & ML Platform)
Senior AI / ML Engineer (MLOps & ML Platform)

Acquism SARL • Amsterdam

On-site
EUR 90,000 - 120,000