Senior Ml Engineer (Token Factory)

Jobgether

Málaga

Presencial

EUR 70.000 - 110.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Jobgether is listing on behalf of a partner company. A Senior Machine Learning Engineer (Token Factory) based in Spain will join a team building large-scale AI systems and high-performance infrastructure for training and serving foundation models at scale.

You will optimize inference, training pipelines, and GPU efficiency on massive GPU fleets, collaborating with ML, systems, and infrastructure experts in a fast-paced research environment.

Formación

  • Strong ML fundamentals, transformer and LLM knowledge.
  • Experience profiling and optimizing GPU workloads with Nsight or PyTorch Profiler.
  • Deep GPU architecture understanding, memory hierarchy and compute vs memory trade-offs.
  • Familiar with attention, KV-cache, Flash Attention, quantization.
  • Experience with distributed training, sharding, and custom kernels.
  • Proficient in Python and modern ML frameworks.
  • Knowledge of CI/CD, version control, and testing.
  • Excellent cross-functional collaboration.

Responsabilidades

  • Drive inference optimization and latency reduction across diverse LLMs.
  • Design and evolve inference engines, including KV-cache and dense/MoE models support.
  • Develop low-precision training/inference pipelines (e.g., FP8) for large GPU clusters.
  • Profile GPU workloads to identify bottlenecks and guide architecture decisions.
  • Collaborate on scalable distributed training/inference systems with sharding and custom kernels.
  • Contribute to engineering best practices: testing, CI/CD, production-grade ML systems.

Conocimientos

ML fundamentals
GPU profiling
GPU architecture
LLM concepts
Distributed training
Python programming
CI/CD
Communication skills
Team collaboration

Herramientas

Nsight
PyTorch Profiler
CUDA

Descripción del empleo

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Engineer (Token Factory) based in Spain.

This role sits at the intersection of large-scale AI systems and high-performance infrastructure, focusing on optimizing how foundation models are trained and served at scale.
You will contribute to a cutting-edge inference and fine-tuning platform designed to push modern LLMs to their performance limits across massive GPU fleets.
The work directly impacts throughput, latency, and cost efficiency for next-generation AI workloads used in production environments.
You will collaborate with highly specialized engineers across ML, systems, and infrastructure domains in a fast-moving, research-driven environment.
The role combines deep ML expertise with systems-level engineering, requiring strong understanding of both model architecture and hardware behavior.
You will help design and improve critical components such as inference engines, training pipelines, and GPU optimization strategies.

Accountabilities:
  • Drive inference optimization efforts by identifying bottlenecks and implementing performance improvements across diverse LLM architectures, improving throughput and reducing latency and cost per token.
  • Contribute to the design and evolution of inference engines, including techniques such as speculative decoding, KV-cache optimization, and support for dense and MoE models.
  • Develop and productionize low-precision training and inference pipelines (e.g., FP8, MXFP4) to maximize efficiency on large GPU clusters.
  • Profile and analyze GPU workloads using modern tooling to identify performance constraints and guide architectural improvements.
  • Collaborate on scalable distributed training and inference systems, including sharding strategies, custom kernels, and hardware-aware optimizations.
  • Contribute to engineering best practices including testing, CI/CD, and maintainable production-grade ML systems.
Requirements:
  • Strong understanding of machine learning fundamentals, particularly transformer architectures and large language models.
  • Hands-on experience profiling and optimizing GPU workloads using tools such as Nsight or PyTorch Profiler.
  • Deep knowledge of GPU architecture, including memory hierarchy and compute vs. memory trade-offs.
  • Familiarity with key LLM concepts such as attention mechanisms, RoPE, KV-cache, Flash Attention, and quantization techniques.
  • Experience with large-scale deep learning training, including distributed systems, sharding strategies, and custom kernel development.
  • Strong software engineering skills, with advanced proficiency in Python and modern ML frameworks.
  • Solid understanding of software engineering practices such as version control, CI/CD pipelines, and unit testing.
  • Strong communication skills with the ability to collaborate effectively in highly technical, cross-functional teams.
  • Competitive compensation package
  • Strong career development and continuous learning opportunities
  • Flexible work environment with high autonomy and ownership
  • Opportunity to work on frontier AI systems at massive scale
  • International, highly skilled, and diverse team environment
How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Applied Ai Engineer
Applied Ai Engineer

Jobgether • Málaga

Híbrido
EUR 60.000 - 100.000
Healthcare coverage
Remote-first / Hybrid-friendly culture
Unlimited PTO
+1
Senior ML Engineer: High-Scale Inference & GPU Optimization
Senior ML Engineer: High-Scale Inference & GPU Optimization

Jobgether • Málaga

Presencial
EUR 70.000 - 110.000
Fullstack Software Engineer - Core
Fullstack Software Engineer - Core

Jobgether • España

Presencial
EUR 50.000 - 80.000
Remote working opportunity within the;
Machine Learning Engineer
Machine Learning Engineer

European Tech Recruit • Madrid

Híbrido
EUR 50.000 - 70.000
Senior Full Stack Engineer - Core UX
Senior Full Stack Engineer - Core UX

Jobgether • España

Presencial
EUR 77.000 - 142.000
Equity options
20 days paid time off
Global team offsites
+3
Deep Learning Engineer (AI/LLM) - Hybrid
Deep Learning Engineer (AI/LLM) - Hybrid

European Tech Recruit • Barcelona

Híbrido
EUR 45.000 - 70.000
AI Full Stack Engineer
AI Full Stack Engineer

European Tech Recruit • Madrid

Presencial
EUR 90.000 - 120.000
Senior AI Engineer (Sant Cugat del Vallès, Spain, Barcelona)
Senior AI Engineer (Sant Cugat del Vallès, Spain, Barcelona)

Biopharma Careers • Sant Cugat del Vallès

Híbrido
EUR 90.000 - 150.000
Flexible working conditions
Life and accident insurance
Health insurance at a competitiveprice
+2
Software Engineer – Ai Systems & Frameworks (Barcelona)
Software Engineer – Ai Systems & Frameworks (Barcelona)

10Xengineers • Pontevedra

Presencial
EUR 42.000 - 60.000
(Senior) AI Engineer - Data & AI Organisation (all genders)
(Senior) AI Engineer - Data & AI Organisation (all genders)

Merck KGaA • Mollet del Vallès

Presencial
EUR 70.000 - 110.000