Staff Engineer (Core & MLOps)

Lever, Inc.

Netherlands

Remote

EUR 120,000 - 150,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Fully remote
Flexible hours
Global collaboration
Open-source exposure

Job summary

Lever, Inc. is seeking a Staff Engineer (Core & MLOps) based in the Netherlands to lead the design and evolution of our core platform, including service contracts, multi-language client libraries, and scalable infrastructure.

You will drive reliability through SLOs and incident reviews, mentor engineers across squads, and collaborate across a globally distributed, remote-first team on cutting-edge data pipelines, AI tooling, and open-source projects.

Qualifications

  • 10+ years building scalable distributed backend systems.
  • Advanced Java with reactive frameworks and strong Python skills.
  • Deep experience with gRPC and Protocol Buffers and schema evolution.
  • Production Kubernetes at scale, Terraform, and Kafka expertise.
  • Experience with telemetry pipelines and reliability engineering.

Responsibilities

  • Architect and evolve core service planes and API contracts.
  • Own the service chassis, client libraries, Helm charts, and deployment pipelines.
  • Lead RFDs on orchestration, multi-cluster routing, and failover.
  • Establish SLOs/SLIs, fault budgets, and incident post-mortems.
  • Mentor engineers across squads and drive engineering best practices.

Skills

Advanced Java
Python
gRPC/Protobuf
Kubernetes
Terraform
Kafka
MLOps
Reliability engineering
Technical writing
Distributed systems

Tools

Temporal
Kubernetes
Terraform
Kafka
SPIRE

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in Netherlands.

Accountabilities:
  • Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health-aware routing, automated canary releases, and operational feedback loops.
  • Own the service chassis and golden path, maintaining and improving multi-language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
  • Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
  • Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real-time billing pipelines, Valkey, and database modernization initiatives.
  • Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi-cluster routing, automated failover, and other critical platform initiatives.
  • Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
  • Participate in shared infrastructure on-call rotations, lead incident post-mortems, and convert operational insights into platform improvements.
  • Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
Requirements:
  • 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
  • Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
  • Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission-critical systems.
  • Hands-on production experience with Kubernetes at scale, Terraform, and event-streaming platforms such as Kafka.
  • Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
  • Strong reliability engineering background, including SLO/SLI definition, blast-radius analysis, fault tolerance, and rigorous service contracts.
  • Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
  • Strong written and interpersonal communication skills suited to a globally distributed, remote-first environment.
  • A curious, continuous-learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
  • Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
  • MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
  • Familiarity with zero-trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
  • Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
  • Experience with large-scale web scraping or crawling, or contributions to distributed-systems and data-extraction open-source projects, is advantageous.
Benefits:
  • Fully remote, remote-first working environment with flexible working hours.
  • Freedom and flexibility to work from the location where you are most productive.
  • Opportunity to work on core infrastructure supporting large-scale web data pipelines and distributed systems.
  • Exposure to cutting-edge open-source technologies, tools, and evolving AI and web data infrastructure.
  • Opportunities to attend conferences and connect with colleagues across the globe.
  • Collaboration with a diverse, multicultural, and globally distributed engineering community.
  • High level of autonomy and organizational trust.
  • Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer - Core & MLOps, Remote Platform Architect
Staff Engineer - Core & MLOps, Remote Platform Architect

Lever, Inc. • Netherlands

Remote
EUR 120,000 - 150,000
Fully remote
Flexible hours
Global collaboration
+1
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • Netherlands

Remote
EUR 90,000 - 130,000
Fully remote work
Engineering Manager - Foundations & Enablement
Engineering Manager - Foundations & Enablement

Lever, Inc. • Netherlands

Remote
EUR 120,000 - 180,000
Monthly remote work stipend
Home office equipment package
Annual international team-building
Senior Backend Engineer: Machine Learning Infrastructure
Senior Backend Engineer: Machine Learning Infrastructure

Lever, Inc. • Netherlands

Remote
EUR 71,000 - 106,000
Fully remote environment
Unlimited vacation
Home-office stipend
+4
MLOps Engineer
MLOps Engineer

EPAM Systems • Netherlands

Hybrid
EUR 85,000 - 120,000
Automation Engineer (Customer Service)
Automation Engineer (Customer Service)

Lever, Inc. • Netherlands

On-site
EUR 70,000 - 110,000
Competitive compensation
Flexible working environment
International team
+4
Head of Platform Engineering
Head of Platform Engineering

Harnham • Randstad

On-site
EUR 120,000 - 180,000
Senior Software Engineer - Java / Kotlin / AI / Full Stack
Senior Software Engineer - Java / Kotlin / AI / Full Stack

Join • Utrecht

On-site
EUR 75,000 - 110,000
Profit sharing
Flexible working hours
Trainer opportunities
+2
Engineering Manager
Engineering Manager

United States Digital Space LLC • Amsterdam

On-site
EUR 110,000 - 160,000
Daily catered lunches
Commuting reimbursement
25 vacation days
+5
Backend Engineers @ Platform team - Massive throughput, strict latency, and high availability r[...]
Backend Engineers @ Platform team - Massive throughput, strict latency, and high availability r[...]

Sprint and Partners • Amsterdam

On-site
EUR 90,000 - 140,000
Top-of-market Amsterdam compensation
Hybrid work
Growth budget
+2