Senior MLOps Engineer: Scale LLM Training & Inference

multiversecomputing

Donostia/San Sebastián

Híbrido

EUR 70.000 - 110.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Indefinite contract
Equal pay guaranteed
Variable performance bonus
Signing bonus
Work visa sponsorship
Relocation package
Private health insurance
Hybrid opportunity
Flexible working hours
Language classes
Discounted meals

Descripción de la vacante

Multiverse Computing is seeking a Senior MLOps Engineer to shape the Training and Inference Optimization team. You will architect infrastructure powering next-gen AI models, bridging systems programming and ML to optimize large-scale training and ultra-high-throughput serving.

You will lead training pipelines, inference orchestration, and lifecycle management while mentoring engineers and driving cost-efficient, production-ready solutions.

Formación

  • 5+ years in MLOps, DevOps, or Software Engineering, with at least 2 years in LLM infrastructure.
  • Expert-level proficiency with PyTorch and the NVIDIA stack (CUDA, NCCL, Triton).
  • Hands-on experience with NVIDIA NeMo (or Megatron‑Bridge) for distributed training and at least two serving options: vLLM, TensorRT-LLM, or SGLang.
  • Experience with SLURM/Flyte/Ray/SkyPilot for cluster management and MLflow for experiment and model management.
  • Deep expertise in Kubernetes and K8s operators (KubeRay, MPI Operator, Run:ai).
  • Mastery of Python and knowledge of C++ or Rust for performance-critical components.
  • Familiarity with high-performance networking (InfiniBand/RoCE) and NVIDIA H200/B200 (Blackwell).

Responsabilidades

  • Training Infrastructure: Architect and maintain scalable distributed training pipelines using NVIDIA NeMo/Nemotron/Megatron‑Bridge.
  • optimise GPU utilisation, manage checkpointing strategies, and implement automated fault tolerance for long-running jobs.
  • Inference Orchestration: Lead the deployment of LLMs using vLLM, TensorRT‑LLM, or SGLang; tune PagedAttention, batching, and quantisation.
  • Workload Orchestration: Use SLURM/Flyte/Ray/SkyPilot to manage ML workloads across clouds and on-prem clusters.
  • Lifecycle Management: Standardise model tracking and versioning with MLflow.
  • Performance Engineering: Profile and optimize across CUDA kernels, NCCL, and Python orchestration.
  • Efficiency & Cost Governance: Optimize cloud/on-prem GPU expenditures with smart scaling.
  • Technical Leadership: Drive roadmap, code reviews, and mentor teams.

Conocimientos

MLOps
DevOps
Software Engineering
LLM infrastructure
PyTorch
CUDA
NCCL
Triton
NVIDIA NeMo
Megatron-Bridge
vLLM
TensorRT-LLM
SGLang
SLURM
Flyte
Ray
SkyPilot
MLflow
Kubernetes
KubeRay
MPI Operator
Run:ai
Python
C++
Rust
InfiniBand
RoCE
NVIDIA H200
NVIDIA B200
Open-source contributions
Model compression
Triton kernels

Herramientas

NVIDIA NeMo
Megatron‑Bridge
vLLM
TensorRT-LLM
SGLang

Descripción del empleo

Multiverse Computing is seeking a Senior MLOps Engineer to shape the Training and Inference Optimization team. You will architect infrastructure powering next-gen AI models, bridging systems programming and ML to optimize large-scale training and ultra-high-throughput serving.

You will lead training pipelines, inference orchestration, and lifecycle management while mentoring engineers and driving cost-efficient, production-ready solutions.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior MLOps Engineer (Training & Inference Optimization)
Senior MLOps Engineer (Training & Inference Optimization)

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 70.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Senior AI Engineer: Scalable AI & MLOps Leader
Senior AI Engineer: Scalable AI & MLOps Leader

UL Solutions • Madrid

Híbrido
EUR 70.000 - 110.000
Annual bonus (target 10%)
ULS University
Two volunteering days per year
+1
MLOps & GenAI Engineer — Scalable AI Platform
MLOps & GenAI Engineer — Scalable AI Platform

Enfuce • España

Presencial
EUR 60.000 - 90.000
Remote work flexibility
Stock option program
Flexible vacation policy
+1
Senior LLM Engineer
Senior LLM Engineer

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 70.000 - 110.000
Indefinite contract.
Equal pay guaranteed.
Variable performance bonus.
+9
AI Platform Engineer for Scalable Inference & Deployment
AI Platform Engineer for Scalable Inference & Deployment

Visa Hunt • España

Presencial
EUR 90.000 - 130.000
Remote-friendly environment
Flexible working options
Competitive compensation package
+1
Senior AI Engineer: Scalable, Production-Ready AI (Remote)
Senior AI Engineer: Scalable, Production-Ready AI (Remote)

UL Solutions • Barcelona

Híbrido
EUR 53.000 - 58.000
Senior MLOps Engineer (ML Workflows Engineering)
Senior MLOps Engineer (ML Workflows Engineering)

JetBrains • Madrid

Presencial
EUR 70.000 - 110.000
Senior MLOps Engineer: Build Scalable ML Pipelines
Senior MLOps Engineer: Build Scalable ML Pipelines

JetBrains • Madrid

Presencial
EUR 70.000 - 110.000
Senior Staff Engineer - Production AI & LLM (Remote)
Senior Staff Engineer - Production AI & LLM (Remote)

Cohere • España

Híbrido
EUR 90.000 - 130.000
Lunch stipend
Health benefits
Parental leave
+5
Senior ML Engineer – LLMOps & Production AI Lead
Senior ML Engineer – LLMOps & Production AI Lead

News Corporation • Bellprat

Híbrido
EUR 55.000 - 70.000
Healthcare plans for you and family
Remote work 3 months/year
Meal benefit via Pluxee
+3