Staff Engineer (Core & MLOps)

Lever, Inc.

España

A distancia

EUR 120.000 - 180.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura completa en un minuto — currículum adaptado y carta de presentación, listos para enviar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Fully remote
Flexible hours
Open-source exposure
Conference attendance
Global collaboration

Descripción de la vacante

Lever, Inc. in Spain is seeking a Staff Engineer (Core & MLOps) to help shape foundational infrastructure powering large-scale web data products.

You’ll own the architecture of core control and context planes that enable AI-driven workflows to operate reliably across multi-cloud environments. The role blends hands-on architecture with technical leadership, mentoring, and cross-functional alignment across teams.

Formación

  • 10+ years of experience building scalable distributed backend systems.
  • Strong Java expertise with reactive frameworks (Vert.x or Netty) and Python proficiency.
  • Deep experience with gRPC and Protocol Buffers, including schema evolution.
  • Hands-on production experience with Kubernetes at scale, Terraform, and Kafka.
  • Experience designing automated telemetry pipelines or feature stores to improve systems.

Responsabilidades

  • Architect and evolve core control and context planes and service contracts.
  • Maintain and improve multi-language client libraries and deployment pipelines.
  • Define inter-service contracts, API gateways, and schema evolution standards.
  • Operate core platform infrastructure across Kubernetes, Terraform, and Confluent Kafka.
  • Lead architectural strategy, RFDs, and reliability engineering practices.
  • Mentor engineers and drive cross-team alignment on platform direction.

Conocimientos

Java
Python
gRPC & Protobuf
Kubernetes
Terraform
Reactive Java
SRE & Reliability
Architecture
Technical Writing
MLOps
Remote Collaboration

Herramientas

Helm
Kafka
Envoy/Istio

Descripción del empleo

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in Spain.

This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you’ll tackle complex distributed systems challenges at production scale. You’ll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands‑on architecture with technical leadership, mentoring, and cross‑functional alignment. In a globally distributed, remote‑first environment, you’ll have significant autonomy to solve challenging infrastructure problems and influence long‑term platform strategy.

Accountabilities:
  • Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health‑aware routing, automated canary releases, and operational feedback loops.
  • Own the service chassis and golden path, maintaining and improving multi‑language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
  • Define and govern inter‑service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
  • Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real‑time billing pipelines, Valkey, and database modernization initiatives.
  • Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi‑cluster routing, automated failover, and other critical platform initiatives.
  • Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
  • Participate in shared infrastructure on‑call rotations, lead incident post‑mortems, and convert operational insights into platform improvements.
  • Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
Requirements:
  • 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
  • Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
  • Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission‑critical systems.
  • Hands‑on production experience with Kubernetes at scale, Terraform, and event‑streaming platforms such as Kafka.
  • Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
  • Strong reliability engineering background, including SLO/SLI definition, blast‑radius analysis, fault tolerance, and rigorous service contracts.
  • Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
  • Strong written and interpersonal communication skills suited to a globally distributed, remote‑first environment.
  • A curious, continuous‑learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
  • Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
  • MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
  • Familiarity with zero‑trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
  • Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
  • Experience with large‑scale web scraping or crawling, or contributions to distributed‑systems and data‑extraction open‑source projects, is advantageous.
Benefits:
  • Fully remote, remote‑first working environment with flexible working hours.
  • Freedom and flexibility to work from the location where you are most productive.
  • Opportunity to work on core infrastructure supporting large‑scale web data pipelines and distributed systems.
  • Exposure to cutting‑edge open‑source technologies, tools, and evolving AI and web data infrastructure.
  • Opportunities to attend conferences and connect with colleagues across the globe.
  • Collaboration with a diverse, multicultural, and globally distributed engineering community.
  • High level of autonomy and organizational trust.
  • Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre‑contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Software Engineer - Reliability, Infrastructure, and Tooling
Senior Software Engineer - Reliability, Infrastructure, and Tooling

Jobgether • España

A distancia
EUR 117.000 - 261.000
Equity participation
Fully remote
Health, dental, and vision
+1
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • España

A distancia
EUR 70.000 - 100.000
Fully remote work environment
Technical leadership opportunities
Software Engineer - Python and K8s
Software Engineer - Python and K8s

Jobgether SRL • España

A distancia
EUR 55.000 - 90.000
Fully distributed remote work
Learning & development budget
Annual compensation review
+7
Staff Backend Engineer - Grafana Second Horizon | Spain | Remote
Staff Backend Engineer - Grafana Second Horizon | Spain | Remote

Grafana • Madrid

Presencial
EUR 94.000 - 113.000
Remote-first team
RSUs
30 days annual leave
+1
Engineering Manager - Foundations & Enablement
Engineering Manager - Foundations & Enablement

Lever, Inc. • España

A distancia
EUR 90.000 - 130.000
Monthly remote work stipend
Home office equipment
Dedicated development fund
+3
Staff Backend Engineer - Grafana Second Horizon | Spain | Remote New Spain (Remote)
Staff Backend Engineer - Grafana Second Horizon | Spain | Remote New Spain (Remote)

Grafana • España

A distancia
EUR 94.025 - 112.830
RSUs
Remote-first culture
Competitive compensation
Senior Backend Engineer - Databases - Analytics | Spain | Remote
Senior Backend Engineer - Databases - Analytics | Spain | Remote

Grafana • España

A distancia
EUR 83.000 - 104.000
Remote work
MLOps Field Engineer
MLOps Field Engineer

Jobgether SRL • España

Presencial
EUR 85.000 - 115.000
Geographically adjusted pay
Annual bonus
Fully distributed work environment
+5
Product Engineer (Platform)
Product Engineer (Platform)

Jobgether • Málaga

Presencial
EUR 103.000 - 155.000
Equity / stock options
AI‑first environment
Flat organizational structure
+2
Senior Data and Analytics Engineer ID89052
Senior Data and Analytics Engineer ID89052

AgileEngine, LLC. • Madrid

Presencial
EUR 70.000 - 90.000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1