Senior ML Engineer: High-Scale Inference & GPU Optimization

Jobgether

Málaga

Presencial

EUR 70.000 - 110.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Jobgether is listing on behalf of a partner company. A Senior Machine Learning Engineer (Token Factory) based in Spain will join a team building large-scale AI systems and high-performance infrastructure for training and serving foundation models at scale.

You will optimize inference, training pipelines, and GPU efficiency on massive GPU fleets, collaborating with ML, systems, and infrastructure experts in a fast-paced research environment.

Formación

  • Strong ML fundamentals, transformer and LLM knowledge.
  • Experience profiling and optimizing GPU workloads with Nsight or PyTorch Profiler.
  • Deep GPU architecture understanding, memory hierarchy and compute vs memory trade-offs.
  • Familiar with attention, KV-cache, Flash Attention, quantization.
  • Experience with distributed training, sharding, and custom kernels.
  • Proficient in Python and modern ML frameworks.
  • Knowledge of CI/CD, version control, and testing.
  • Excellent cross-functional collaboration.

Responsabilidades

  • Drive inference optimization and latency reduction across diverse LLMs.
  • Design and evolve inference engines, including KV-cache and dense/MoE models support.
  • Develop low-precision training/inference pipelines (e.g., FP8) for large GPU clusters.
  • Profile GPU workloads to identify bottlenecks and guide architecture decisions.
  • Collaborate on scalable distributed training/inference systems with sharding and custom kernels.
  • Contribute to engineering best practices: testing, CI/CD, production-grade ML systems.

Conocimientos

ML fundamentals
GPU profiling
GPU architecture
LLM concepts
Distributed training
Python programming
CI/CD
Communication skills
Team collaboration

Herramientas

Nsight
PyTorch Profiler
CUDA

Descripción del empleo

Jobgether is listing on behalf of a partner company. A Senior Machine Learning Engineer (Token Factory) based in Spain will join a team building large-scale AI systems and high-performance infrastructure for training and serving foundation models at scale.

You will optimize inference, training pipelines, and GPU efficiency on massive GPU fleets, collaborating with ML, systems, and infrastructure experts in a fast-paced research environment.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Ml Engineer (Token Factory)
Senior Ml Engineer (Token Factory)

Jobgether • Málaga

Presencial
EUR 70.000 - 110.000
Senior AI Systems Engineer — Large-Scale Training (Hybrid)
Senior AI Systems Engineer — Large-Scale Training (Hybrid)

SGI • Barcelona

Híbrido
EUR 90.000 - 130.000
Equity
Relocation support
Hybrid work
+1
Senior ML Ops Engineer - Remote GPU Inference
Senior ML Ops Engineer - Remote GPU Inference

Pragmatike • Madrid

Presencial
EUR 90.000 - 130.000
Senior ML Platform Engineer - Scale & GenAI
Senior ML Platform Engineer - Scale & GenAI

Preply • Barcelona

Presencial
EUR 80.000 - 110.000
Generous learning allowance
Health insurance
Access to mental health support
Fullstack AI Platform Engineer — Data, ML & UI
Fullstack AI Platform Engineer — Data, ML & UI

Jobgether • España

Presencial
EUR 50.000 - 80.000
Remote working opportunity within the;
Remote ML Ops Engineer — Scalable GPU Inference
Remote ML Ops Engineer — Scalable GPU Inference

Pragmatike • Madrid

A distancia
EUR 90.000 - 120.000
Global MLOps Field Engineer: AI Infra & Cloud Architect
Global MLOps Field Engineer: AI Infra & Cloud Architect

Jobgether • España

Híbrido
EUR 65.000 - 95.000
Learning budget USD 2,000 per year
40 days of annual leave
Travel opportunities
+2
Senior Applied Research Engineer | Barcelona, Spain, Hybrid
Senior Applied Research Engineer | Barcelona, Spain, Hybrid

SGI • Barcelona

Híbrido
EUR 90.000 - 130.000
Equity
Relocation support
Hybrid work
+1
ML Architect: Next-Gen AI Architectures
ML Architect: Next-Gen AI Architectures

Jobgether • España

Presencial
EUR 70.000 - 110.000
Competitive equity
Global distributed team
Production deployment focus
Senior ML Engineer - Scalable AI Infra & Pipelines
Senior ML Engineer - Scalable AI Infra & Pipelines

IFS • Madrid

Presencial
EUR 45.000 - 65.000