Senior ML Engineer (Evaluation)

kaiko.ai

Zürich

Hybrid

CHF 140.000 - 190.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Competitive salary
Pension plan
25 vacation days
Learning and development budget
Commuting subsidy
Team offsites

Zusammenfassung

kaiko.ai is hiring a Senior Evaluation ML Engineer to own the evaluation stack for scalable, production‑grade ML pipelines. The role emphasizes observability, reliability, and close collaboration with ML researchers and product to translate clinical requirements into dependable eval signals.

You will work from Zurich or Amsterdam, spending roughly half your time in the office, and contribute to a growing, international team developing clinical AI tools for diagnostics and care coordination.

Qualifikationen

  • Excellent Python skills and strong software engineering fundamentals: testing, modular design, CI/CD, code review, and monorepo tooling.
  • Proven experience building and operating ML inference services or MLOps infrastructure at scale, ideally for large language or multimodal models.
  • Hands-on experience with distributed compute and GPU workloads; familiarity with Ray, CUDA toolchains, and container runtimes (Docker/Kubernetes).

Aufgaben

  • Own AI factory orchestration for evaluation workloads, with Dagster as the primary orchestration layer.
  • Maintain and evolve inference services powering evaluation runs, including resource management and throughput.
  • Ensure the functional integrity of the eval stack through testing and validation across configurations.
  • Own Eval/MLOps end-to-end: deployments, model registry, artifact versioning, rollout/rollback procedures, and observability.
  • Develop towards a technical lead: set engineering direction and support other engineers.

Kenntnisse

Python
CI/CD
Monorepo tooling
Testing
Distributed compute
GPU workloads
MLOps
Evaluation services
Technical leadership

Tools

Dagster
Ray
Docker
Kubernetes
TensorRT-LLM
Triton Inference Server
vLLM
OpenAI Evals / eval frameworks
Terraform / IaC

Jobbeschreibung

kaiko building a next-generation agentic clinical AI assistant that helps clinicians reason across patient data, guidelines, and diagnostics.

Healthcare decisions are rarely made by a single person or from a single data source. kaiko’s assistant maintains longitudinal patient context across encounters, clinicians, and institutions, enabling collaboration, second opinions, and complex diagnostic workflows. The system is designed to operate safely in real clinical environments, with human oversight, auditability, and regulatory alignment at its core.

Our assistant core supports broadly applicable clinical tasks such as patient data navigation, guideline interaction, multimodal interaction (chat and voice), and care coordination. On top of this foundation, we are developing specialized diagnostic agents in areas such as oncology, radiology, and pathology.

We build in close collaboration with leading hospitals and research centers, including the Netherlands Cancer Institute (NKI). kaiko is a well-funded company with a growing international team, operating from Zurich and Amsterdam.

About the Role

Kaiko’s Multimodal Large Language Model (MLLM) is trained on domain-specific, high-complexity medical data. Reaching clinical-grade performance demands a comprehensive evaluation stack that is fast, reliable, and deeply integrated with our model development loop.

As a Senior Evaluation ML Engineer, you will own the engineering stack to run evaluations at scale, from efficient inference across a growing set of frontier models to async evaluation against a wide array of clinical benchmarks, enabling automated orchestration of our pipelines with a strong eye for observability and production-grade system organisation. You will work closely with other ML researchers and product to translate research and clinical requirements into reliable and well-engineered eval signals.

As a Senior ML Evaluation Engineer you will

  • Own AI factory orchestration for evaluation workloads, with Dagster as the primary orchestration layer: design, operate, and mature the pipelines and workflows that run large-scale evaluation jobs, and extend automation across the stack wherever possible.
  • Maintain and evolve the inference services that power evaluation runs, including cluster- and actor-level resource management, ensuring correctness, reproducibility, and throughput as the model and benchmark zoo grows.
  • Ensure the functional integrity of the eval stack through rigorous testing and validation: verify model integrations, confirm expected behaviour across configurations, and support ML researchers in understanding model outputs.
  • Own Eval/MLOps end-to-end: service deployments, model registry and artifact versioning, eval database organisation, rollout and rollback procedures, and post-deployment observability.
  • Develop towards a technical lead: set engineering direction, make architectural decisions, and support other engineers in execution.

You will be based in Zurich or Amsterdam, with the expectation of spending ∼50% of your time in the office.

About you

  • Excellent Python skills and strong software engineering fundamentals: testing, modular design, CI/CD, code review, and monorepo tooling.
  • Proven experience building and operating ML inference services or MLOps infrastructure at scale, ideally for large language or multimodal models.
  • Hands‑on experience with distributed compute and GPU workloads: familiarity with frameworks such as Ray, CUDA toolchains, and container runtimes (Docker/Kubernetes or equivalent).
  • Experience with model serving frameworks such as vLLM, TensorRT‑LLM, Triton Inference Server, or similar.
  • Experience with workflow orchestration tools, with a preference for Dagster; ability to design reliable, maintainable pipeline DAGs.
  • Familiarity with the full deployment lifecycle, from containerisation and config management to observability, alerting, and incident response.
  • Ability to read and reason about model internals at a low level: tokenisation, numerical precision, tensor shapes, and inference‑time behaviour.
  • Prior experience in the medical domain is not required, but a strong motivation to push the frontier of clinical AI through excellent engineering is.

Nice to have

  • Experience acting as a technical lead: setting direction on an engineering sub‑system, making architectural trade‑offs, and guiding other engineers.
  • Familiarity with eval frameworks (lm‑eval‑harvest, OpenAI Evals, HF Evaluate) and benchmark integration pipelines.
  • Background in software‑defined infrastructure, IaC tooling (Terraform, Pulumi), or cloud‑native deployments on AWS/GCP/Azure.
  • Safety and reliability engineering mindset: experience with red‑teaming, load testing, or quality practices for production AI systems.

We are excited to gather a broad range of perspectives in our team, as we believe it will help us build better products to support a broader set of people. If you’re excited about us but don’t fit every single qualification, we still encourage you to apply: we’ve had incredible team members join us who didn’t check every box.

At kaiko, we believe the best ideas come from collaboration, ownership and ambition. We’ve built a team of international experts where your work has direct impact. Here’s what we value:

  • Ownership: You’ll have the autonomy to set your own goals, make critical decisions, and see the direct impact of your work.
  • Collaboration: You’ll have to approach disagreement with curiosity, build on common ground and create solutions together.
  • Ambition: You’ll be surrounded by people who set high standards for themselves and others, who see obstacles as opportunities, and who are relentless in their work to create better outcomes for patients.

In addition, we offer

  • An attractive and competitive salary, a good pension plan and 25 vacation days per year.
  • Great offsites and team events to strengthen the team and celebrate successes together.
  • A EUR 1,000 learning and development budget to help you grow.
  • Autonomy to do your work the way that works best for you, whether you have a kid or prefer early mornings.
  • An annual commuting subsidy.

Our interview process is designed to assess mutual fit across skills, motivation, and values. It typically includes the following steps:

  • Screening call: A short conversation to align on your motivation, career goals, and initial fit for the role.
  • Technical interview: A deep dive into your problem‑solving approach through a technical challenge, case study, or role‑specific scenario.
  • Onsite meeting (optional): You’ll meet team members across functions to explore collaboration dynamics, team fit, and day‑to‑day context.
  • Final executive conversation: A discussion with a member of the executive team focused on long‑term alignment, cultural fit, and shared expectations for impact.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

ML Platform Engineer
ML Platform Engineer

kaiko.ai • Zürich

Hybrid
CHF 120.000 - 180.000
Competitive salary
Pension plan
25 vacation days
+3
Technical Project Manager
Technical Project Manager

kaiko.ai • Zürich

Hybrid
CHF 110.000 - 160.000
Salary & pension
Vacation days
Team events
+3
Site Reliability Engineer
Site Reliability Engineer

kaiko.ai • Zürich

Vor Ort
CHF 140.000 - 190.000
Good pension plan
25 vacation days per year
Offsites and team events
+2
Senior Evaluation ML Engineer for Clinical AI Platform
Senior Evaluation ML Engineer for Clinical AI Platform

kaiko.ai • Zürich

Hybrid
CHF 140.000 - 190.000
Competitive salary
Pension plan
25 vacation days
+3
Senior Full Stack Engineer (f/m/d)
Senior Full Stack Engineer (f/m/d)

Embodied AI • Schweiz

Hybrid
CHF 94.000 - 169.000
Stock options
Hybrid/Remote CET±4
Mental health app access
+1
Founding AI Engineer
Founding AI Engineer

Ahead Health • Zürich

Vor Ort
CHF 145.000 - 155.000
Relocation assistance negotiable
Direct access to medical experts
Significant equity
Postdoctoral Researcher in Multimodal Reasoning Models for Oncology
Postdoctoral Researcher in Multimodal Reasoning Models for Oncology

ETH Zürich • Basel

Vor Ort
CHF 70.000 - 90.000
Access to unique multimodal clinical datasets
Competitive salary
Excellent research infrastructure
Postdoctoral Researcher in Multimodal Reasoning Models for Oncology 100%
Postdoctoral Researcher in Multimodal Reasoning Models for Oncology 100%

ETH Zürich • Basel

Vor Ort
CHF 80.000 - 100.000
Access to unique multimodal clinical datasets
Competitive salary
Collaboration with leading experts
+1
Postdoctoral Researcher in Multimodal Reasoning Models for Oncology
Postdoctoral Researcher in Multimodal Reasoning Models for Oncology

Immigration Policy Lab • Basel

Vor Ort
CHF 80.000 - 100.000
Competitive salary
Access to unique multimodal clinical datasets
Excellent research infrastructure
+1
Senior ML Engineer
Senior ML Engineer

Jobgether • Schweiz

Vor Ort
CHF 140.000 - 210.000
Competitive compensation adjusted to 3
Equity where applicable
Mentorship and knowledge sharing