Senior ML Engineer, Inference & Optimization

Nebius Group

Zürich

Vor Ort

CHF 150.000 - 220.000

Vollzeit

Vor 10 Tagen
Bewerbungsgenerator

Verschicke keinen 08/15-Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Competitive compensation
Career growth and learning机会
Flexibility and ownership
Collaborative and innovative culture
Opportunity to work on impactful AI项目
International environment and talented

Zusammenfassung

Nebius Group in Zürich seeks a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts through production deployment. You will work on model internals, inference engines, serving architecture, and benchmarking to improve latency, throughput, memory efficiency, GPU utilization and cost per token while maintaining model quality.

This is a hands‑on role that collaborates with kernel and platform engineers to diagnose bottlenecks, compare serving configurations, and

Qualifikationen

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM/VLM systems.
  • Knowledge of modern inference stacks such as vLLM, TensorRT-LLM, Triton.

Aufgaben

  • Own optimization work for model families, endpoints, or serving backends.
  • Run engine comparisons and recommend serving configurations.
  • Debug quality or performance regressions during production rollouts.
  • Deploy, configure, benchmark, and extend inference engines (e.g., vLLM, Triton).
  • Build reproducible benchmarks for latency, throughput, and cost per token.

Kenntnisse

Python
PyTorch
LLM deployment
Performance optimization
Inference systems

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
NVIDIA Dynamo
Ray Serve
KServe

Jobbeschreibung

Nebius Group in Zürich seeks a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts through production deployment. You will work on model internals, inference engines, serving architecture, and benchmarking to improve latency, throughput, memory efficiency, GPU utilization and cost per token while maintaining model quality.

This is a hands‑on role that collaborates with kernel and platform engineers to diagnose bottlenecks, compare serving configurations, and

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior AI Inference Engineer — Kernel & Edge Optimization
Senior AI Inference Engineer — Kernel & Edge Optimization

Lever, Inc. • Schweiz

Remote
EUR 137.000 - 222.000
Remote-first team
International team
Cutting-edge AI research
+3
Senior Backend Engineer — Low-Latency AI Search
Senior Backend Engineer — Low-Latency AI Search

AI Chopping Block • Zürich

Vor Ort
CHF 140.000 - 200.000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Applied ML Engineer — Agentic Search Platform
Senior Applied ML Engineer — Agentic Search Platform

AI Chopping Block • Zürich

Vor Ort
CHF 120.000 - 180.000
Competitive compensation
Career growth and learning
International environment
Senior Applied ML Engineer – Agent-Native Search
Senior Applied ML Engineer – Agent-Native Search

Carbon Data Solutions • Zürich

Vor Ort
CHF 112.000 - 169.000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Systems Engineer - Low-Latency AI Search Platform
Senior Systems Engineer - Low-Latency AI Search Platform

Nebius • Zürich

Vor Ort
CHF 140.000 - 200.000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Applied ML Engineer - Agentic Search Platform
Senior Applied ML Engineer - Agentic Search Platform

Nebius Group • Zürich

Vor Ort
CHF 170.000 - 260.000
Competitive compensation
Career growth
Flexibility and ownership
+3
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Schweiz

Remote
EUR 137.000 - 222.000
Remote-first team
International team
Cutting-edge AI research
+3
Machine Learning Engineer - Production AI Systems (Zurich)
Machine Learning Engineer - Production AI Systems (Zurich)

Embodied AI • Zürich

Hybrid
CHF 120.000 - 165.000
Senior HPC & AI Network Architect for Scalable AI Infra
Senior HPC & AI Network Architect for Scalable AI Infra

NVIDIA • Zürich

Vor Ort
CHF 180.000 - 240.000
Senior Data Scientist: ML, Data Eng & Growth (Hybrid)
Senior Data Scientist: ML, Data Eng & Growth (Hybrid)

NZZ-Mediengruppe • Zürich

Vor Ort
CHF 150.000 - 190.000
Flexible working hours
Hybrid work-from-home
Birthday off
+1