Senior Platform Engineer — AI Infrastructure

Helloprint

Rotterdam

On-site

EUR 110,000 - 160,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Room to grow
International team
Real impact
Ownership from day one
HelloBenefits: gym & meals

Job summary

HelloPrint seeks an experienced Production SRE / Platform Engineer to own reliability, observability, and safe deployment across high-velocity microservices. You will drive SLOs, automation, and cost governance, while partnering with product teams to ensure scalable AI runtime infrastructure and robust cloud operations.

Join an international team in Rotterdam/Valencia, shaping the future of a rapidly growing AI-driven platform with pragmatic guardrails and efficient workflows.

Qualifications

  • Proven experience operating high-traffic production systems with uptime and safe release cycles.
  • Troubleshooting Linux, containers, databases, Redis, and cloud networking.
  • Hands-on with Google Cloud Platform and Terraform.
  • Cost governance and FinOps mindset for cloud/AI runtimes.
  • Experience configuring Sentry and Cloud Monitoring for SLIs.
  • Programming with Python or TypeScript/JavaScript; Laravel ecosystem familiarity.
  • AI operations familiarity with LLMs, embeddings, and rate limits.
  • Engineering mindset focused on durable guardrails and automation.

Responsibilities

  • Define, track, and enforce SLOs/SLIs and error budgets across core services.
  • Expand Telemetry, Sentry, and Google Cloud Monitoring across microservices.
  • Evolve CI/CD pipelines with canary deployments and automated health gates.
  • Architect and monitor runtime infrastructure for AI agents and pipelines.
  • Drive FinOps and cost governance across Google Cloud workloads.
  • Lead on-call incident response and post-mortems; derive automated tests.
  • Forecast capacity; plan disaster recovery validations with RTO/RPO targets.
  • Own IaC and security using Terraform and IAM least privilege.

Skills

SRE/Platform
Distributed tracing
GCP
Terraform
FinOps
Python/TypeScript
Laravel experience
DevEx tooling

Tools

Sentry

Job description

HelloPrint is mid-transformation. The entire platform is being rebuilt from the ground up: new frontend, new pricing engine, new content engine, new product engine. Everything agent-ready. We ship in a week what used to take a year. What we do not yet have is someone who owns the reliability of all of it: the SLOs, the cost, the AI runtime, and the guardrails that let product engineers move fast without breaking things. That is this role.

Core Objective: Take full technical ownership of production reliability, distributed observability, deployment safety, cost optimization, and AI runtime infrastructure across high-velocity microservices and cloud workloads.
What you will do:
  • SLOs & Error Budgets: Define, track, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error-budget policies across core customer journeys and critical services (including checkout, payments, catalog pipelines, and routing engines).

  • Distributed Observability & Telemetry: Expand Telemetry, Sentry + Google Cloud Monitoring distributed tracing, and automated diagnostic tooling across microservices, Laravel Horizon queue workers on Redis, and Google Cloud infrastructure.

  • Deployment Safety & Progressive Delivery: Partner with product teams to evolve CI/CD pipelines with canary traffic shifting, automated SLO-driven rollbacks, and automated health gates in GitHub Actions.

  • AI Platform Operations & FinOps: Architect, monitor, and scale the runtime infrastructure supporting AI agents, semantic pipelines, and background automation. Take full ownership of runtime cost control, model and token budget tracking, latency profiles, rate limits, queue backpressure, and provider availability.

  • Cloud Infrastructure Cost & Budget Optimization: Partner with engineering leadership to drive continuous FinOps practices across Google Cloud workloads, actively identifying resource inefficiencies, optimizing compute/storage footprint, and enforcing infrastructure budget guardrails.

  • Incident Management & Post-Mortems: Lead on-call incident response and blameless post-mortems, systematically turning root causes into automated tests, synthetic checks, and architectural guardrails.

  • Resilience, Capacity & DR: Drive capacity forecasting, dependency isolation, automated load testing, and disaster recovery validations against strict RTO/RPO targets.

  • Infrastructure as Code & Security: Own declarative infrastructure workflows using Terraform and Google Cloud Run, ensuring strict IAM least privilege, Secret Manager, and deterministic environments.

  • Toil Elimination & DevEx: Build pragmatic internal tooling, runbooks, and self-service deployment primitives that eliminate firefighting and enable product engineers to ship reliably.

What we are looking for:
  • Production SRE & Platform Background: Proven experience operating and scaling high-traffic distributed production systems where uptime, low latency, and safe release cycles are paramount.

  • Deep Diagnostic Capabilities: Strong troubleshooting skills across Linux environments, containerized runtimes, relational/NoSQL databases, Redis queues, and cloud network boundaries.

  • GCP & Serverless Container Stacks: Extensive hands-on experience with Google Cloud Platform, Google Cloud Run, and infrastructure automation via Terraform.

  • Cost Governance & FinOps Mindset: Demonstrated ability to monitor, analyze, and optimize cloud and AI runtime costs, striking the right balance between performance, reliability, and budgetary efficiency.

  • Targeted Monitoring & Tracing: Practical experience configuring Sentry (for error reporting and distributed tracing across Laravel and Astro frontends) alongside Google Cloud Monitoring for SLI tracking and canary gates.

  • Automation & Scripting Skills: Proficiency in Python or TypeScript/JavaScript for platform tooling, combined with practical working familiarity with modern PHP in a Laravel ecosystem.

  • AI Operations Affinity: Familiarity with operationalizing LLM integrations, embeddings/vector workflows, rate-limited external APIs, or background task orchestration.

  • Engineering Mindset: Pragmatic builder with high technical agency who values durable system guardrails and deterministic automation over manual interventions.

What we offer:
  • Room to grow: opportunities develop fast, within your role and across HelloPrint

  • International environment: work with 100+ professionals from 20+ nationalities in Rotterdam or Valencia

  • Real impact: shape the future of a high-growth, AI-driven company where your ideas actually land

  • Ownership from day one: work with experienced leaders, take initiative, and drive projects with freedom to innovate

  • Hello-Benefits: including 24/7 HelloFit gym access, UrbanSportsClub discount, the best events and company-sponsored healthy meals every day

At HelloPrint, we believe diversity makes us better. We’re building a workplace where everyone can be themselves, grow, and do their best work regardless of background, identity, or perspective. We welcome applicants from all backgrounds.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE & AI Platform Operations)
Senior Site Reliability Engineer (SRE & AI Platform Operations)

Helloprint • Rotterdam

On-site
EUR 90,000 - 130,000
HelloFit gym access
UrbanSportsClub discount
Company-sponsored healthy meals
Senior AI Engineer
Senior AI Engineer

Helloprint • Rotterdam

On-site
EUR 80,000 - 140,000
AI Engineer
AI Engineer

Helloprint • Rotterdam

On-site
EUR 80,000 - 140,000
Salary €80k-€140k
Room to grow
Global team
+1
Retention Growth Builder
Retention Growth Builder

Helloprint • Rotterdam

On-site
EUR 50,000 - 70,000
24/7 HelloFit gym access
UrbanSportsClub discount
Company-sponsored healthy meals
Senior Test Automation Engineer
Senior Test Automation Engineer

Helloprint • Rotterdam

On-site
EUR 50,000 - 67,000
Gym access
Sports discount
Company events
+1
Senior Business & Performance Analyst
Senior Business & Performance Analyst

HelloPrint • Rotterdam

On-site
EUR 70,000 - 90,000
Competitive Salary
24/7 access to company gym
Company-sponsored healthy meals
+2
AI Infra Reliability Lead — Platform Engineer
AI Infra Reliability Lead — Platform Engineer

Helloprint • Rotterdam

On-site
EUR 110,000 - 160,000
Room to grow
International team
Real impact
+2
DevOps/Platform Engineer
DevOps/Platform Engineer

Wellis • Rotterdam

Hybrid
EUR 61,000 - 89,000
Salary range €5,500 – €8,000 gross per
Pension plan
Learning budget
+7
Engineering Manager
Engineering Manager

United States Digital Space LLC • Amsterdam

Hybrid
EUR 110,000 - 160,000
Daily catered lunches
Commuting reimbursement
25 vacation days
+5
Experienced Back-end & Platform Engineer
Experienced Back-end & Platform Engineer

Clockworks • Rotterdam

On-site
EUR 55,000 - 75,000
Competitive compensation
Flexible working conditions
Support for professional ambitions