Staff Engineer (Core & MLOps)

Lever, Inc.

Ireland

Remote

EUR 150,000 - 190,000

Full time

30 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Fully remote
Remote-first
Autonomy
Global collaboration

Job summary

Lever, Inc. in Ireland is seeking a Staff Engineer (Core & MLOps) to shape foundational infrastructure powering large-scale web data products and distributed teams. You’ll own the architecture of core control and context planes enabling AI-driven workflows at production scale.

The role blends hands-on architecture with technical leadership, mentoring, and creating reliable platform services. This remote-first position offers significant autonomy and impact across multiple squads.

Qualifications

  • 10+ years creating scalable distributed backend systems.
  • Advanced Java with reactive frameworks (Vert.x/Netty).
  • Deep experience with gRPC and Protocol Buffers.
  • Production Kubernetes at scale and Terraform usage.

Responsibilities

  • Architect and evolve control and context planes with SLO enforcement and canary releases.
  • Maintain multi-language Java and Python client libraries and deployment pipelines.
  • Define inter-service contracts including gRPC, Protocol Buffers, and API gateway strategies.
  • Operate core platform infra across Kubernetes, Terraform, and Kafka ecosystems.
  • Lead architecture discussions via RFDs on workflow and gateway orchestration.
  • Mentor engineers across squads and drive reliable software practices.

Skills

Java
Python
gRPC
Kubernetes
Terraform
Kafka
Reactive

Tools

Confluent Kafka
Istio

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in Ireland.

This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you’ll tackle complex distributed systems challenges at production scale. You’ll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands‑on architecture with technical leadership, mentoring, and cross‑functional alignment. In a globally distributed, remote‑first environment, you’ll have significant autonomy to solve challenging infrastructure problems and influence long‑term platform strategy.

Accountabilities
  • Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health‑aware routing, automated canary releases, and operational feedback loops.
  • Own the service chassis and golden path, maintaining and improving multi‑language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
  • Define and govern inter‑service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
  • Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real‑time billing pipelines, Valkey, and database modernization initiatives.
  • Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi‑cluster routing, automated failover, and other critical platform initiatives.
  • Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
  • Participate in shared infrastructure on‑call rotations, lead incident post‑mortems, and convert operational insights into platform improvements.
  • Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
Requirements
  • 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
  • Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
  • Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission‑critical systems.
  • Hands‑on production experience with Kubernetes at scale, Terraform, and event‑streaming platforms such as Kafka.
  • Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
  • Strong reliability engineering background, including SLO/SLI definition, blast‑radius analysis, fault tolerance, and rigorous service contracts.
  • Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
  • Strong written and interpersonal communication skills suited to a globally distributed, remote‑first environment.
  • A curious, continuous‑learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
  • Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
  • MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
  • Familiarity with zero‑trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
  • Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
  • Experience with large‑scale web scraping or crawling, or contributions to distributed‑systems and data‑extraction open‑source projects, is advantageous.
Benefits
  • Fully remote, remote‑first working environment with flexible working hours.
  • Freedom and flexibility to work from the location where you are most productive.
  • Opportunity to work on core infrastructure supporting large‑scale web data pipelines and distributed systems.
  • Exposure to cutting‑edge open‑source technologies, tools, and evolving AI and web data infrastructure.
  • Opportunities to attend conferences and connect with colleagues across the globe.
  • Collaboration with a diverse, multicultural, and globally distributed engineering community.
  • High level of autonomy and organizational trust.
  • Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.
How Jobgether works

We use an AI‑powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role’s core requirements. Our system identifies the top‑fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre‑contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Manager - Foundations & Enablement
Engineering Manager - Foundations & Enablement

Lever, Inc. • Ireland

Remote
EUR 110,000 - 150,000
Remote stipend
Home office kit
Co-working access
Applied ML Engineer
Applied ML Engineer

Jobgether SRL • Ireland

On-site
EUR 70,000 - 120,000
Location: Ireland
Staff Security Operations Engineer
Staff Security Operations Engineer

Jobgether • Ireland

On-site
EUR 120,000 - 180,000
Annual learning & development budget:
Remote-first work environment
In-person team sprints
+4
Onboarding Operations Specialist
Onboarding Operations Specialist

Jobgether SRL • Ireland

Remote
EUR 25,000 - 56,000
Fully remote
Flexible PTO
Flexible hours
+7
Senior Full Stack AI-first Saas Engineer
Senior Full Stack AI-first Saas Engineer

Jobgether • Ireland

Remote
EUR 70,000 - 90,000
Fully remote work environment
Premium medical insurance
18 days annual paid vacation
+2
Senior Product Analyst
Senior Product Analyst

Jobgether SRL • Ireland

Remote
EUR 75,000 - 110,000
Fully remote
Vacation days
Wellness days
+5
Tech Architect Practitioner
Tech Architect Practitioner

Jobgether SRL • Ireland

On-site
EUR 120,000 - 180,000
Direct executive exposure
Diverse client projects
Entrepreneurial consulting environment
+1
Senior Site Reliability Engineer / Kubernetes
Senior Site Reliability Engineer / Kubernetes

Jobgether • Ireland

Remote
EUR 90,000 - 130,000
100% remote within EU time zones
Flexible working hours
Ownership and impact in a fast-paced,技
+1
Senior Partner Sales Manager - Global System Integrator (GSI)
Senior Partner Sales Manager - Global System Integrator (GSI)

Jobgether • Ireland

On-site
EUR 90,000 - 150,000
Globally distributed work environment
Team sprints twice yearly
USD 2,000 learning budget per year
+9
Senior Security Engineer, Offensive Security
Senior Security Engineer, Offensive Security

Jobgether • Ireland

On-site
EUR 119,000 - 170,000
Fully remote
EU compensation
Equity
+1