Site Reliability Engineer

Yuno

Amsterdam

On-site

EUR 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work
Home office bonus
Work equipment
Stock options
Health plan
Flexible days off
Growth courses

Job summary

Yuno is seeking a Staff Site Reliability Engineer to set the technical direction for reliability across the platform that provisions, deploys, and manages AI agents at scale on AWS for payments across 190+ countries. This is a fully remote, full-time role for an experienced engineer ready to own the reliability strategy and drive architectural evolution.

You will lead the design of event-driven communication, define SLOs/error budgets, and mentor teams to improve observability and resilience

Qualifications

  • Experience designing and owning event-driven architectures and messaging systems (e.g., Kafka, NATS, RabbitMQ)
  • Deep AWS knowledge (EC2, VPC, IAM, S3, RDS) with strong networking basics
  • Infrastructure as Code with Terraform or Pulumi
  • Kubernetes and Docker in production environments
  • Observability and SLO/SLI management with monitoring tools like Datadog
  • Chaos engineering and resilience testing across production systems
  • Strong debugging of distributed async flows and automationcoding
  • Proven technical leadership setting reliability standards across teams
  • Excellent verbal and written English communication

Responsibilities

  • Define and drive reliability strategy and incident practices at scale
  • Architect and evolve platform reliability, messaging, and event-driven patterns
  • Own cloud infrastructure, IaC automation, and scalable deployment
  • Build observability dashboards, tracing, and alerting for platform health
  • Lead incident response, postmortems, and root-cause analyses across teams
  • Mentor senior and mid-level engineers to raise the reliability bar
  • Champion chaos experiments and resilience patterns to prevent production incidents

Skills

Event-driven architecture
AWS
Observability & SLOs
Chaos engineering
Distributed systems debugging
Leadership
English proficiency

Tools

Kafka/NATS/RabbitMQ
Terraform/Pulumi
Kubernetes
Docker
Datadog

Job description

Remote · Full Time · Individual Contributor · +7 Years of Experience

Site Reliability Engineer
Who We Are

Yuno is the AI-native operating system of global commerce, powering the financial infrastructure of enterprise merchants, banks, and wallets. Through a single API, Yuno connects them to pay-ins, payouts, fraud prevention, KYC/KYB, and stablecoins globally, so they can operate everywhere. Agnostic by design and connected to 1,000+ payment methods and 460+ integrations in 190+ countries, Yuno optimizes acceptance rates, reduces costs, and strengthens security through specialized AI agents that learn from every transaction. Global brands including McDonald's, NetEase Games, GoFundMe, and Rappi run their payments on Yuno.

About The Role

Yuno is looking for a Staff Site Reliability Engineer to set the technical direction for reliability across our infrastructure — starting with the platform that provisions, deploys, and manages AI agents at scale on AWS, the system powering payments across 190+ countries. The platform is in production and growing, and we need the most senior reliability voice in the room to evolve the architecture and make sure it stays reliable, observable, and ready to scale.

This is not a "maintain what exists" role, and it's not a single-system role. You'll own the reliability strategy — driving architectural decisions, designing event-driven communication, defining how we measure and defend reliability, and setting the standards other engineering teams build on.

How AI Shows Up in This Role
  • The platform you own is Yuno's AI agent infrastructure — provisioning and deploying AI agents at scale, plus the agents that route payments and prevent fraud. Keeping the AI-native layer reliable is the core of the role

  • AI is our default execution layer: you're encouraged to use AI-assisted tooling across automation, runbooks, incident analysis, and root-cause investigations, and to help define how the wider engineering org adopts it. We care how you use it, not whether you do

Your Contribution Will Be
  • Reliability strategy and standards — define the SLO culture, error-budget policy, and incident practices that scale across engineering teams, turning reliability from firefighting into a measurable, org-wide discipline

  • Platform architecture and evolution — drive architectural decisions as the platform matures; the deciding voice on choosing technologies, designing systems, and when to evolve the infrastructure

  • Messaging and event-driven architecture — design and own the messaging layer for inter-service communication, replacing synchronous patterns with durable, reliable async messaging

  • Infrastructure and deployment — own the cloud infrastructure, automate provisioning with IaC, and ensure the platform scales reliably as transaction volume grows

  • Observability — build the monitoring, tracing, and alerting that keeps the platform healthy; when something breaks at 3am, your dashboards and alerts should explain why before anyone has to dig

  • Incident leadership and mentorship — act as the senior escalation point for the hardest production problems, run blameless postmortems and root-cause analyses that turn into permanent fixes, and raise the reliability bar by mentoring senior and mid-level engineers

  • Chaos engineering mindset — continuous fault injection and resilience experiments that surface weaknesses before they turn into incidents, plus identifying and proposing resilience patterns to prevent those failures from reaching production.

What Success Looks Like

Within your first 6–12 months, you've set the reliability strategy for the platform, driven at least one major architectural evolution (event-driven messaging, streaming reliability, or observability), and engineering teams have adopted the SLO and error-budget framework you defined. You're the person Yuno trusts with the hardest reliability calls.

Skills You Need
Minimum Qualifications
  • Event-driven architecture and messaging systems — you've designed and owned systems around message queues (Kafka, NATS, RabbitMQ) and understand at-least-once delivery, consumer groups, dead letters, and backpressure; you've migrated a system from synchronous to async

  • Deep AWS — EC2, VPC, IAM, S3, and RDS — with strong networking fundamentals, since inter-service communication runs over the internal VPC

  • Infrastructure as Code — Terraform or Pulumi, reviewed in PRs rather than clicked in consoles

  • Kubernetes and Docker in production — container lifecycle, resource limits, health checks, and orchestration at scale

  • Observability and SLOs — Datadog fluency or equivalent (dashboards, monitors, APM, distributed tracing), and a track record defining and operating SLOs, SLIs, and error budgets across services

  • Chaos engineering and resilience testing — hands‑on experience with fault injection, game days, or chaos experiments (Gremlin, Chaos Mesh, AWS FIS, or similar) to harden production systems

  • Distributed systems debugging — you've diagnosed async flows and cascading failures in production and can explain what broke and how you fixed it; comfortable coding for automation and tooling (Go, Python, or similar)

  • Databases — solid SQL (PostgreSQL) and NoSQL (MongoDB, Redis): when to use each, indexing, replication, and performance tuning

  • Proven technical leadership — you've set reliability standards, influenced architecture across teams, and mentored engineers, not just owned your own scope

  • English — advanced proficiency, written and spoken

Preferred Qualifications
  • AI / MLOps infrastructure — running AI workloads in production (model serving, LLM inference, GPU/resource management, and agent evaluation/observability tools like LangFuse, LangSmith, Braintrust, or MLflow)

  • Multi‑tenant container platforms — running customer or user workloads in containers (Replit, Railway, Fly.io, or internal PaaS)

  • Data pipelines and orchestration — Airflow, Prefect, or similar; data warehouses like Databricks, Snowflake, or BigQuery a plus

  • Incident management and on‑call tooling — PagerDuty, Opsgenie, or incident.io

  • Experience in the payments industry

Nice to Have
  • ECS experience

  • s6-overlay for container process supervision

  • Experience with AI agent framework ecosystems

  • Spanish proficiency

What We Offer at Yuno
  • Competitive Compensation

  • Remote Work — you can work from everywhere

  • Home Office Bonus — a one‑time allowance to set up your ideal home office

  • Work Equipment

  • Stock Options

  • Health Plan wherever you are

  • Flexible Days Off

  • Language, Professional, and Personal Growth courses

Disclaimer

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed or wish to exercise your data protection rights, please contact us at hiring@y.uno.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Yuno Totalmente remoto · Mundial hace 2 días
Site Reliability Engineer Yuno Totalmente remoto · Mundial hace 2 días

Tamarind Intelligence • Amsterdam

Hybrid
EUR 140,000 - 210,000
Remote Work
Health Plan wherever you are
Stock Options
+2
Senior Corporate IT Engineer Yuno Totalmente remoto · Mundial hace 2 días
Senior Corporate IT Engineer Yuno Totalmente remoto · Mundial hace 2 días

Tamarind Intelligence • Amsterdam

Hybrid
EUR 70,000 - 110,000
Remote work
Home Office Bonus
Work Equipment
+4
Staff Engineer — Data Platform
Staff Engineer — Data Platform

Yuno • Netherlands

On-site
EUR 70,000 - 100,000
Competitive Compensation
Remote Work
Home Office Bonus
+5
Head of Product - Core Team
Head of Product - Core Team

Yuno • Amsterdam

Remote
EUR 100,000 - 150,000
Remote work everywhere
Home Office Bonus
Health Plan
+1
Engineering Manager – Data Platform
Engineering Manager – Data Platform

Yuno • Netherlands

On-site
EUR 75,000 - 100,000
Competitive Compensation
Remote Work
Home Office Bonus
+3
Remote Site Reliability Engineer for AI Payments Platform
Remote Site Reliability Engineer for AI Payments Platform

Yuno • Amsterdam

Remote
EUR 140,000 - 200,000
Remote work
Home office bonus
Work equipment
+4
Staff Site Reliability Engineer - Remote, AI-Driven Platform
Staff Site Reliability Engineer - Remote, AI-Driven Platform

Tamarind Intelligence • Amsterdam

Hybrid
EUR 140,000 - 210,000
Remote Work
Health Plan wherever you are
Stock Options
+2
Staff Engineer, AI-Driven Payments Platform (Remote)
Staff Engineer, AI-Driven Payments Platform (Remote)

Yuno • Netherlands

On-site
EUR 120,000 - 150,000
Competitive compensation
Remote work
Home office allowance
+5
Remote Engineering Manager, AI-Driven Data Platform
Remote Engineering Manager, AI-Driven Data Platform

Yuno • Amsterdam

On-site
EUR 120,000 - 180,000
Remote work
Home office bonus
Stock options
+4
Remote Senior Corporate IT Engineer — AI-Driven Automation
Remote Senior Corporate IT Engineer — AI-Driven Automation

Yuno • Amsterdam

On-site
EUR 86,000 - 121,000
Competitive Compensation
Remote Work — you can work from everyh
Home Office Bonus — setup allowance
+5