Senior Observability Infrastructure Engineer

United States Digital Space LLC

Amsterdam

On-site

EUR 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

the company in Amsterdam is seeking a Senior Observability Infrastructure Engineer to join our Platform Engineering team. You will help design, build, and run the logging, metrics, and tracing platform at scale across on-premise and Kubernetes environments.

You will own the hybrid infrastructure, automate operations with Go or Python, and optimize the data pipelines to support hundreds of product teams. This role emphasizes reliability, security, and self-service through automated guardrails.

Qualifications

  • 10+ years of experience in observability or related platform/infrastructure domains.
  • Hands-on experience operating core telemetry data stores at scale (logging, metrics, tracing).
  • Production Kubernetes experience on-premises and/or cloud with kubectl proficiency.
  • Proficient in Go or Python and building automation tooling and IaC mindset.
  • Experience tuning large-scale distributed tracing and logging pipelines.

Responsibilities

  • Design and implement the future architecture of the logging and metrics systems across global regions.
  • Own hybrid infrastructure, managing the lifecycle of 1,500+ servers on bare metal and Kubernetes.
  • Automate operations with Go or Python; improve CI pipelines for safe, automated cluster changes.
  • Optimize Elasticsearch, Prometheus, VictoriaMetrics, and OpenTelemetry for peak performance.
  • Participate in on-call rotations while improving self-service guardrails and safe API access.

Skills

Observability
Linux
Kubernetes
Go
Python

Tools

Elasticsearch/OpenSearch
VictoriaMetrics
Grafana Tempo
Prometheus
Clickhouse

Job description

This is the company

the company provides payments, data, and financial products in a single solution for customers like Facebook, Uber, H&M, and Microsoft - making us the financial technology platform of choice. At the company, everything we do is engineered for ambition.

For our teams, we create an environment with opportunities for our people to succeed, backed by the culture and support to ensure they are enabled to truly own their careers. The people of the company are motivated individuals who tackle unique technical challenges at scale and solve them as a team. Together, we deliver innovative and ethical solutions that help businesses achieve their ambitions faster.

Senior Observability Infrastructure Engineer

We are looking for an experienced Observability Infrastructure Engineer to join our Platform Engineering organization. You will be part of the team responsible for building and running Observability pillars on premise and on Kubernetes. Our systems collect, process, and store the logs, metrics, and traces that allow hundreds of product teams to monitor their services in real time.

This is a role for a builder and a problem solver who enjoys deep technical troubleshooting across distributed systems and then turns recurring issues into automated, repeatable solutions. You will work in a large-scale environment where we manage petabytes of data and thousands of servers. We are currently in the middle of a major transformation: focusing on automation of operations and enabling self service for our users.

What you will do
  • Build the next generation of our platform: Design and implement the future architecture of our logging and metrics systems. You will play a key role in redesigning our infrastructure to support new global regions, ensuring data isolation and regulatory compliance in different geographies, and more.
  • Own infrastructure operations: You will take full ownership of our hybrid infrastructure, managing the lifecycle of over 1,500 servers across both bare-metal and Kubernetes environments.
  • Automate to reduce toil: You will write code in Go or Python to eliminate manual operational tasks. Your goal is to build self-healing systems that do not require manual intervention during the night. You will improve our CI pipelines to ensure that changes to our clusters are safe, predictable, and automated.
  • Optimize for scale and performance: You will dive deep into performance bottlenecks within our distributed tracing and logging pipelines. We deal with high-volume data streams that can overwhelm standard configurations. You will tune our Elasticsearch clusters, optimize Prometheus and VictoriaMetrics storage, and ensure our OpenTelemetry implementation can handle peak traffic without missing a beat.
  • Reliability and Engineering: You will participate in on-call rotations, but your primary focus will be engineering solutions that stop alerts from firing in the first place. You will help us upgrade our stack to the latest versions and ensure our platform remains secure and performant. You will improve the self-service experience by implementing automated guardrails and quota management to prevent noisy tenants from destabilizing the platform, while designing safer API access patterns for our users.
What you bring
  • 10+ years of experience in the observability domain or in a relevant platform/infrastructure domain.
  • Observability Stack Expertise: You have hands-off experience operating core telemetry data stores at scale e.g. Elasticsearch/Opensearch/VictoriaLogs/Clickhouse for logging, Prometheus/VictoriaMetrics for metrics and Grafana Tempo for distributed tracing.
  • Linux Experience: You understand the operating system at a kernel level and can debug complex networking, file system, and performance issues on both bare metal and virtualized hardware .
  • Production Kubernetes Experience: Proven hands-on experience operating, and troubleshooting production workloads on Kubernetes (on-prem and/or cloud), including strong day-to-day use of kubectl and Kubernetes primitives (e.g. Namespaces, Pods, Deployments/StatefulSets, Services, Ingress, ConfigMaps/Secrets)
  • Software Engineering Mindset: You are proficient in Go or Python and do not just write scripts; you build tools and automation platforms that treat infrastructure as code.
Nice to have
  • Experience with large scale, multi tenant isolation and quota or cost governance approaches for telemetry platforms.
  • Familiarity with regulated environments where security, audibility, and data handling requirements shape platform design decisions.
Our Diversity, Equity, and Inclusion commitments

Our unique approach is a product of our diverse perspectives. This diversity of backgrounds and cultures is essential in helping us maintain our momentum. Our business and technical challenges are unique, and we need as many different voices as possible to join us in solving them - voices like yours. No matter who you are or where you're from, we welcome you to be your true self at the company.

This role is based out of our Amsterdam office. We are an office-first company and value in-person collaboration; we do not offer remote-only roles.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - Observability
Staff Engineer - Observability

United States Digital Space LLC • Amsterdam

On-site
EUR 150,000 - 185,000
Senior Observability Platform Engineer - On-Prem & Kubernetes
Senior Observability Platform Engineer - On-Prem & Kubernetes

United States Digital Space LLC • Amsterdam

On-site
EUR 140,000 - 180,000
Staff Software Engineer - Observability
Staff Software Engineer - Observability

EngineersOfAI • Amsterdam

On-site
EUR 120,000 - 190,000
Infrastructure Developer
Infrastructure Developer

United States Digital Space LLC • Amsterdam

On-site
EUR 90,000 - 130,000
Staff Software Engineer - Observability
Staff Software Engineer - Observability

Adyen • Amsterdam

On-site
EUR 130,000 - 180,000
Staff Engineer - Observability
Staff Engineer - Observability

Adyen • Amsterdam

On-site
EUR 180,000 - 240,000
Infrastructure Developer (Go)
Infrastructure Developer (Go)

United States Digital Space LLC • Amsterdam

On-site
EUR 90,000 - 130,000
Staff Engineer - Observability
Staff Engineer - Observability

EngineersOfAI • Amsterdam

Hybrid
EUR 140,000 - 190,000
Senior Platform Engineer
Senior Platform Engineer

XpertDirect • Amsterdam

On-site
EUR 70,000 - 90,000
Staff Software Engineer
Staff Software Engineer

Super • Amsterdam

On-site
EUR 110,000 - 170,000
Medical / Health Insurance
Open Annual Leave
Employee Assistance Programme
+1