Senior MLOps Engineer (Training & Inference Optimization)

multiversecomputing

Donostia/San Sebastián

Presencial

EUR 70.000 - 110.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Indefinite contract
Equal pay guaranteed
Variable performance bonus
Signing bonus
Work visa sponsorship
Relocation package
Private health insurance
Hybrid opportunity
Flexible working hours
Language classes
Discounted meals

Descripción de la vacante

Multiverse Computing is seeking a Senior MLOps Engineer to shape the Training and Inference Optimization team. You will architect infrastructure powering next-gen AI models, bridging systems programming and ML to optimize large-scale training and ultra-high-throughput serving.

You will lead training pipelines, inference orchestration, and lifecycle management while mentoring engineers and driving cost-efficient, production-ready solutions.

Formación

  • 5+ years in MLOps, DevOps, or Software Engineering, with at least 2 years in LLM infrastructure.
  • Expert-level proficiency with PyTorch and the NVIDIA stack (CUDA, NCCL, Triton).
  • Hands-on experience with NVIDIA NeMo (or Megatron‑Bridge) for distributed training and at least two serving options: vLLM, TensorRT-LLM, or SGLang.
  • Experience with SLURM/Flyte/Ray/SkyPilot for cluster management and MLflow for experiment and model management.
  • Deep expertise in Kubernetes and K8s operators (KubeRay, MPI Operator, Run:ai).
  • Mastery of Python and knowledge of C++ or Rust for performance-critical components.
  • Familiarity with high-performance networking (InfiniBand/RoCE) and NVIDIA H200/B200 (Blackwell).

Responsabilidades

  • Training Infrastructure: Architect and maintain scalable distributed training pipelines using NVIDIA NeMo/Nemotron/Megatron‑Bridge.
  • optimise GPU utilisation, manage checkpointing strategies, and implement automated fault tolerance for long-running jobs.
  • Inference Orchestration: Lead the deployment of LLMs using vLLM, TensorRT‑LLM, or SGLang; tune PagedAttention, batching, and quantisation.
  • Workload Orchestration: Use SLURM/Flyte/Ray/SkyPilot to manage ML workloads across clouds and on-prem clusters.
  • Lifecycle Management: Standardise model tracking and versioning with MLflow.
  • Performance Engineering: Profile and optimize across CUDA kernels, NCCL, and Python orchestration.
  • Efficiency & Cost Governance: Optimize cloud/on-prem GPU expenditures with smart scaling.
  • Technical Leadership: Drive roadmap, code reviews, and mentor teams.

Conocimientos

MLOps
DevOps
Software Engineering
LLM infrastructure
PyTorch
CUDA
NCCL
Triton
NVIDIA NeMo
Megatron-Bridge
vLLM
TensorRT-LLM
SGLang
SLURM
Flyte
Ray
SkyPilot
MLflow
Kubernetes
KubeRay
MPI Operator
Run:ai
Python
C++
Rust
InfiniBand
RoCE
NVIDIA H200
NVIDIA B200
Open-source contributions
Model compression
Triton kernels

Herramientas

NVIDIA NeMo
Megatron‑Bridge
vLLM
TensorRT-LLM
SGLang

Descripción del empleo

Multiverse Computing

Multiverse is a well-funded, fast-growing deep-tech company founded in 2019. We are the largest quantum software company in the EU and have been recognized by CB Insights (2023 and 2025) as one of the 100 most promising AI companies in the world. With 180+ employees and growing, our team is fully multicultural and international. We deliver hyper-efficient software for companies seeking a competitive edge through quantum computing and artificial intelligence. Our flagship products, CompactifAI and Singularity, address critical needs across various industries:

  • CompactifAI is a groundbreaking compression tool for foundational AI models based on Tensor Networks. It enables the compression of large AI systems—such as language models—to make them significantly more efficient and portable.
  • Singularity is a quantum- and quantum-inspired optimization platform used by blue-chip companies to solve complex problems in finance, energy, manufacturing, and beyond. It integrates seamlessly with existing systems and delivers immediate performance gains on classical and quantum hardware.

You’ll be working alongside world‑leading experts to develop solutions that tackle real‑world challenges. We’re looking for passionate individuals eager to grow in an ethics‑driven environment that values sustainability and diversity. We’re committed to building a truly inclusive culture—come and join us.

About the Role

We are seeking a Senior MLOps Engineer to steer the technical vision of our Training and Inference Optimization team. In this high‑impact role, you will architect the infrastructure that powers our next‑generation AI models. You will bridge the gap between systems programming and machine learning, optimizing large‑scale LLM training via NVIDIA NeMo and building ultra‑high‑throughput serving systems using vLLM, TensorRT-LLM, and SGLang. Your mission is to ensure our models are not only state‑of‑the‑art but also production‑hardened, cost‑efficient, and performant at scale.

Key Responsibilities
  • Training Infrastructure: Architect and maintain scalable distributed training pipelines using NVIDIA NeMo/Nemotron/Megatron‑Bridge. You will optimise GPU utilisation, manage complex checkpointing strategies, and implement automated fault tolerance for long‑running jobs.
  • Inference Orchestration: Lead the deployment of LLMs using vLLM, TensorRT‑LLM, or SGLang. You will implement and tune cutting‑edge techniques—including PagedAttention, continuous batching, and advanced quantisation (AWQ/FP8)—to maximise throughput and minimise TPOT (Time Per Output Token).
  • Workload Orchestration: Utilize SLURM/Flyte/Ray/SkyPilot to manage and scale ML workloads across diverse cloud providers and on‑prem clusters, ensuring seamless resource shifting and cost‑effective execution.
  • Lifecycle Management: Standardise model tracking, versioning, and transition workflows using MLflow (or similar tool), ensuring reproducible training runs and a clear path from research to production.
  • Performance Engineering: Conduct deep‑dive profiling and bottleneck analysis across the full stack—from CUDA kernels and NCCL collective communications to Python‑level orchestration.
  • Efficiency & Cost Governance: Monitor and optimise cloud and on‑prem GPU expenditures through intelligent scaling policies and high‑density resource packing.
  • Technical Leadership: Set the bar for engineering excellence. You will drive the roadmap, perform rigorous code reviews, and mentor junior and mid‑level engineers.
Required Qualifications
  • Experience: 5+ years in MLOps, DevOps, or Software Engineering, with a minimum of 2 years dedicated to LLM infrastructure.
  • Deep Learning Ecosystem: Expert‑level proficiency with PyTorch and the NVIDIA stack (CUDA, NCCL, Triton).
  • Specialised Tooling: Hands‑on experience with NVIDIA NeMo (or Megatron‑Bridge) for distributed training and at least two of the following for serving: vLLM, TensorRT‑LLM, or SGLang.
  • Orchestration & Lifecycle: Proven experience with SLURM/Flyte/Ray/SkyPilot for cluster management and MLflow (or similar tool) for experiment and model management.
  • Infrastructure: Deep expertise in Kubernetes and K8s operators (e.g., KubeRay, MPI Operator, or Run:ai).
  • Systems Programming: Mastery of Python and a functional understanding of C++ or Rust for performance‑critical components.
  • Next‑Gen Hardware: Familiarity with high‑performance networking (InfiniBand/RoCE) and NVIDIA H200/B200 (Blackwell) architectures.
Preferred Skills
  • Active contributions to relevant open‑source projects (vLLM, SGLang, SkyPilot, or NeMo).
  • Proven track record with model compression (Sparsity, Distillation, or Quantisation).
  • Experience writing or optimising custom Triton kernels.

Expertise in ML observability stacks (Prometheus, Grafana, Jaeger).

Perks & Benefits
  • Indefinite contract
  • Equal pay guaranteed.
  • Variable performance bonus.
  • Signing bonus.
  • We offer work visa sponsorship (If applicable).
  • Relocation package (if applicable).
  • Private health insurance.
  • Eligibility for educational budget according to internal policy.
  • Hybrid opportunity.
  • Flexible working hours.
  • Language classes and discounted lunch options.
  • Working in a high paced environment, working on cutting edge technologies.
  • Career plan. Opportunity to learn and teach.

As an equal opportunity employer, Multiverse Computing is committed to building an inclusive workplace. The company welcomes people from all different backgrounds, including age, citizenship, ethnic and racial origins, gender identities, individuals with disabilities, marital status, religions and ideologies, and sexual orientations to apply.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior LLM Engineer
Senior LLM Engineer

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 70.000 - 110.000
Indefinite contract.
Equal pay guaranteed.
Variable performance bonus.
+9
Engineering Manager, AI / ML
Engineering Manager, AI / ML

Multiverse Computing • Donostia/San Sebastián

Presencial
EUR 70.000 - 90.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+4
Engineering Manager, AI / ML
Engineering Manager, AI / ML

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 70.000 - 90.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+3
Machine Learning Engineer Intern
Machine Learning Engineer Intern

Multiverse Computing • Donostia/San Sebastián

Híbrido
EUR 11.000 - 19.000
6-month internship contract
Hybrid opportunity
Relocation package (if applicable)
+2
Engineering Manager, AI / ML | Relocation Offered
Engineering Manager, AI / ML | Relocation Offered

Multiverse Computing • Donostia/San Sebastián

Híbrido
EUR 90.000 - 120.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Research AI Director
Research AI Director

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 95.000 - 120.000
Indefinite contract
Variable performance bonus
Private health insurance
+2
Technical Engineering Director - AI / LLMs
Technical Engineering Director - AI / LLMs

Multiverse Computing • Donostia/San Sebastián

Híbrido
EUR 120.000 - 180.000
Indefinite contract
Private health insurance
Hybrid opportunity
+2
Senior Software Engineer
Senior Software Engineer

Multiverse Computing • España

Híbrido
EUR 70.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Senior Data Engineer
Senior Data Engineer

Quantumsoftware • Donostia/San Sebastián

Híbrido
EUR 70.000 - 110.000
Indefinite contract
Equal pay
Signing bonus
+6
Senior Data Engineer
Senior Data Engineer

Quantum Software • Donostia/San Sebastián

Híbrido
EUR 70.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+7