Senior Software Engineer - LLM Ops & Evals

DataSnipper

Amsterdam

On-site

EUR 120,000 - 170,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

DataSnipper in Amsterdam is hiring a Senior Software Engineer for LLM Ops & Eval. You will own the LLM gateway, manage multi-provider inference, and ensure reliable, auditable AI workflows at scale.

You’ll build a platform that supports versioned datasets and evaluation runs used by product and ML teams. You will collaborate across engineering to deploy models across clouds, maintain observability, and implement security controls.

Qualifications

  • 5+ years in backend or platform engineering with production Python
  • Experience building or running LLM inference infrastructure: gateway, multi-provider, failover, cost attribution
  • Experience with cloud infrastructure and IaC across Azure and GCP using Terraform
  • Experience running a shared service in production: on-call, incidents, postmortems, SLOs
  • Hands-on experience with observability tools and dashboards
  • Privacy and compliance work: PII handling, anonymisation, retention, access control
  • Experience building platform/shared-service capabilities for internal teams

Responsibilities

  • Own the LLM gateway: routing, failover, rate limits, retries, cost attribution
  • Deploy, version and deprecate models across clouds and regions
  • Build the shared evaluation platform: versioned datasets, experiments, run schemas, reporting
  • Own infrastructure for async, long-running AI workloads
  • Define and maintain integration contracts between platform and teams
  • Collaborate with product and ML engineers to deliver self-service platform capabilities

Skills

Backend engineering
LLM inference infra
Cloud IaC (Terraform)
Production on-call & SLOs
Observability (OpenTelemetry, Grafana)
Privacy & compliance
Platform/shared services

Tools

Terraform
OpenTelemetry
Grafana

Job description

Which one is yours?

Time to find the right role for you at DataSnipper

The time is now

DataSnipper is a place where you can dive into an idea and bring your unique touch to work. How do we do it? We don't just settle for the way things are, we listen to each other, we look out for each other, we sit down to see how we could improve together. So, we're always looking for people who think differently.

Open positions
Senior Software Engineer - LLM Ops & Evals
Location

Amsterdam

Employment Type

Full time

Location Type

Hybrid

Department

Engineering

Every AI call in DataSnipper goes through us. Product teams do not talk to model providers directly, they talk to our gateway. We are responsible for how inference is routed, how it fails over, what it costs, and how anyone can tell whether the output is any good.

The second half of the job is evaluation. We are building the platform teams use to measure AI quality: versioned datasets, experiment tracking, and evaluation runs they can act on. It is a hard problem and largely an open one, so you will have real influence over how we solve it.

This is a small team with a large blast radius. You will own real production systems, set the standards other teams build against, and see your work in front of hundreds of thousands of users in audit and finance.

About DataSnipper

Audit and finance are still massively manual and we are changing that. DataSnipper is a $1B, bootstrapped unicorn with 600,000+ users across 180+ countries, already embedded in the daily workflows of top audit and accounting firms.
Now, we are taking things further with our Excel Agent, bringing AI directly into where the work actually happens. Unlike generic AI tools, we do not sit on the sidelines. Our AI operates inside Excel, with access to real documents and audit evidence, meaning it does not just generate answers, it does the work, with full traceability.
We are not just applying AI, we are redefining how audit gets done. If you want to build something category-defining at scale, this is the place.

What you will do
Technical Delivery
  • Own the LLM gateway: routing, provider failover, rate limits, retries, and cost attribution across multiple model providers

  • Deploy, version and deprecate models across clouds, regions and environments, including quota and capacity planning, managed as infrastructure as code

  • Build the shared evaluation platform: versioned datasets, experiment tracking, run and result schemas, reporting, and trace linkage back to the run

  • Own the infrastructure for async and long-running AI workloads

Reliability, Security & On-Call
  • Own observability for AI traffic: latency, retries and fallbacks, token usage, cost and errors, per team and per use case

  • Take part in the on-call rotation, run incidents, and close the follow-ups

  • Implement the security and compliance controls the platform is held to: retention, access control, RBAC and SSO

Collaboration & Impact
  • Define and maintain clean integration contracts between the platform and the teams that consume it

  • Partner with product and ML engineers to turn their requirements into platform capabilities that are self-service rather than a request queue

What you will bring
Must-Have
  • 5+ years in backend or platform engineering, with strong production Python

  • Experience building or running LLM inference infrastructure: a gateway or routing layer with multiple providers, failover, rate limiting and cost attribution

  • Experience with cloud at the infrastructure level and infrastructure as code (we run across Azure and GCP with Terraform)

  • Experience running a shared service in production: on-call, incidents, postmortems, SLOs

  • Hands-on experience with observability tools (OpenTelemetry, Grafana), including instrumenting services and designing dashboards and alerts

  • Comfort with privacy and compliance work: PII handling, anonymisation, retention, access control

  • Experience building platform or shared-service capabilities consumed by multiple internal teams

Nice-to-Have
  • Experience with LLM or agent evaluation

  • Temporal or another durable workflow engine

  • Self-hosted inference, capacity planning, load testing

  • Synthetic data or document anonymisation pipelines

  • Document AI: VLMs, OCR, structured extraction and the metrics that go with it

  • Domain experience in audit, accounting or fintech

What we expect
  • Ownership: You own work end-to-end, anticipate issues, and ensure high-quality delivery without close supervision

  • Growth Mindset: You encourage open feedback exchange and provide clear, balanced feedback that helps others grow

  • Collaboration: You build strong cross-functional relationships and influence peers through expertise, data, and empathy

  • Adaptability: You navigate ambiguity calmly, model positive behavior, and help peers adjust through clear communication

  • Judgment: You exercise sound judgment in ambiguous situations, balance speed and accuracy, and adjust priorities proactively

Recruitment steps
  • Recruiter screen

  • Hiring Manager interview

  • Peer programming session

  • System design interview

  • Final interviews with Engineering leadership

Ready to join?

Blaze a new path into your career and take the next step

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

DataSnipper • Amsterdam

On-site
EUR 120,000 - 160,000
Hybrid work (Amsterdam-based)
28 vacation days
Excellent salary
+6
Senior Applied AI Engineer – Agent Runtime
Senior Applied AI Engineer – Agent Runtime

DataSnipper • Amsterdam

On-site
EUR 90,000 - 140,000
Hybrid work (Amsterdam-based)
28 vacation days
Excellent salary
+6
Senior Technical Product Manager — Document Intelligence
Senior Technical Product Manager — Document Intelligence

DataSnipper • Amsterdam

Hybrid
EUR 120,000 - 180,000
Equity
Pension
28 vacation days
+7
Senior Site Reliability Engineer
Senior Site Reliability Engineer

DataSnipper • Amsterdam

Hybrid
EUR 90,000 - 140,000
Equity
Pension
Vacation days
+7
Senior Applied AI Engineer – Agent Runtime (Alwin by DataSnipper)
Senior Applied AI Engineer – Agent Runtime (Alwin by DataSnipper)

DataSnipper • Amsterdam

Hybrid
EUR 110,000 - 160,000
Hybrid work (Amsterdam-based)
Pension plan
Stock participation plan
+4
Forward Deployed Engineer
Forward Deployed Engineer

DataSnipper • Amsterdam

On-site
EUR 90,000 - 130,000
Hybrid work (Amsterdam-based)
28 vacation days
Excellent salary
+5
Site Reliability Engineer
Site Reliability Engineer

DataSnipper • Amsterdam

Hybrid
EUR 90,000 - 120,000
Equity
Pension plan
Hybrid work
+8
Solutions Architect
Solutions Architect

Datasnipper • Amsterdam

On-site
EUR 120,000 - 160,000
Equity
Excellent salary package
Pension plan
+8
Senior Technical Product Manager — Agent Runtime
Senior Technical Product Manager — Agent Runtime

DataSnipper • Amsterdam

Hybrid
EUR 110,000 - 150,000
Equity (Stock Appreciation Rights)
Pension plan
Hybrid work
+3
Customer Success Manager - DACH
Customer Success Manager - DACH

DataSnipper • Amsterdam

Hybrid
EUR 70,000 - 100,000
Equity Shares
Pension Plan
28 Vacation Days
+8