ml engineer for LLM inference

HireHi

España

Presencial

EUR 120.000 - 180.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca la empresa.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Bonuses up to $5000 for referrals
Wellness days: 7
Vacation: 28 days
Health insurance allowance: up to $100
Training reimbursement 50%
English lessons discount
Office equipment provided

Descripción de la vacante

Social Discovery Group ищет эксперта по большим языковым моделям для ускорения инференса и масштабирования моделей в продакшн. Вы будете работать над распределённой инференцией на многопроцессорных кластерах, улучшать системы тонкой настройки и обучения LLM, а также вести команды NLP и CV, формируя дорожную карту исследований.

Требуется глубокий практический опыт в оптимизации инференса, знание PyTorch и Transformer-библиотек, а также владение английским или русским языками на продвинутом

Formación

  • Глубокий практический опыт оптимизации вывода LLM в продакшн с SGLang, vLLM или TensorRT-LLM.
  • Опыт распределённой инференции или обучения крупных моделей, включая MoE, параллелизм и multi-node GPU.
  • Понимание производительности инференса: KV cache, ядра внимания, батчинг, квантование и профилирование GPU.
  • Опыт обучения и тонкой настройки LLM, включая пост-обучение (RLHF/DPO).
  • Лидерство в техническом плане: руководство инженерами, код и код-ревью.

Responsabilidades

  • Ускорять и масштабировать вывод LLM в продакшн с использованием SGLang, KV и префиксного кэширования.
  • Запуск распределённой инференции для очень крупных моделей на многопроцессорных кластерах на нескольких узлах.
  • Бенчмарк новых GPU-серверов и адаптация к ним сервисного кода.
  • Техническое руководство NLP и CV команд: ревью экспериментов, направление исследований.
  • Обучение и финетюнинг языковых моделей, улучшение хитроумных системы агентов и чат-алгоритмов.
  • Отслеживание исследований и open-source; формирование ML-дорожной карты.
  • Сотрудничество с валидационными, контент- и датасет-отделами для проектирования экспериментов.

Conocimientos

LLM inference optimization
Distributed inference
MLOps / deployment
PyTorch
Transformers
English or Russian
Leadership

Herramientas

SGLang
vLLM
TensorRT-LLM
CUDA
Triton

Descripción del empleo

Описание

Social Discovery Group is a group of social discovery companies that addresses loneliness, isolation, and disconnection by transforming virtual intimacy into the new normal. Its products are social entertainment platforms designed to connect people online across different cultures and regions.

Задачи
  • Speed up and scale LLM inference in production using SGLang, KV and prefix caching, batching, quantization, and speculative decoding
  • Run distributed inference for very large models with up to 1T+ parameters across multi-GPU and multi-node setups
  • Benchmark new GPU servers and hardware, bring them into production, and adapt serving code to them
  • Lead the NLP and CV teams technically by reviewing experiments, setting direction, and stepping in early when needed
  • Train and fine-tune language models, and improve the agent harnesses and chat algorithm that run on them
  • Track cutting-edge research and open-source work in inference and post-training, and turn it into the ML roadmap
  • Collaborate with validation, content, and dataset preparation teams to design experiments and measure model quality
Требования
  • Deep hands-on experience optimizing LLM inference in production with SGLang, vLLM, or TensorRT-LLM
  • Experience with distributed inference or training of large models, including MoE, tensor/expert/pipeline parallelism, and multi-node GPU clusters
  • Strong understanding of inference performance, including KV cache, attention kernels, batching, quantization, and GPU profiling
  • Experience training and fine-tuning LLMs, including post-training such as RLHF or DPO
  • Proven technical leadership, including guiding engineers through reviews, mentoring, and technical decisions while continuing to write code
  • Proficiency with PyTorch, transformers, and related libraries
  • Advanced English or Russian
  • Будет плюсом: experience at AI-focused startups or companies such as Character AI or OpenAI, backend engineering experience with Python, Go, or C#, knowledge of scalable deployment systems, CUDA or Triton kernel development, a computer vision background, experience accelerating large generative image or video models, experience with multimodal LLMs, first-author papers or notable open-source work such as contributions to SGLang, vLLM, or post-training libraries, a degree in CS, math, or physics from a strong program (MSc or PhD)
Условия
  • The initial pay level or pay range for this role will be shared with candidates during the recruitment process and before the commencement of employment
  • 28 Calendar days of vacation per year
  • 7 Wellness days per year that can be used to deal with household issues or recover without taking sick leave
  • Bonuses up to $5000 for recommending successful applicants for positions in the company
  • 50% Payment for professional training, international conferences, and meetings
  • Corporate discount for English lessons
  • Health benefits: if not eligible for corporate medical insurance, employees receive compensation of up to $1,000 gross per year for health insurance or doctors’ fees for themselves and close relatives
  • The company provides equipped workplaces and necessary equipment in its offices or co-working locations; elsewhere, it reimburses workplace costs up to $1,000 gross once every 3 years
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

ML инженер
ML инженер

Enfint • Barcelona

Presencial
EUR 60.000 - 90.000
Medical Scheme
Commuting Allowance
Life Insurance
+7
ml engineer for renewable energy asset monitoring
ml engineer for renewable energy asset monitoring

Enfint • Barcelona

Presencial
EUR 60.000 - 120.000
Medical Scheme
Commuting Allowance
Life Insurance
+4
ml engineer in gaming technology
ml engineer in gaming technology

Enfint • Barcelona

Presencial
EUR 60.000 - 90.000
Robust benefits package
Global career opportunities
ml engineer for social casino games
ml engineer for social casino games

HireHi • Barcelona

Presencial
EUR 70.000 - 120.000
ml engineer for gaming products
ml engineer for gaming products

Enfint • Barcelona

Presencial
EUR 70.000 - 100.000
Benefits package
International opportunities
No travel
+1
machine learning engineer for LLM solutions
machine learning engineer for LLM solutions

Enfint • Barcelona

Presencial
EUR 70.000 - 120.000
Комплексные планы медицинского страхия
30 дней оплачиваемого отпуска
Удалённая работа 3 месяца в год
+5
Senior LLM Engineer
Senior LLM Engineer

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 70.000 - 110.000
Indefinite contract.
Equal pay guaranteed.
Variable performance bonus.
+9
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Migx • Barcelona

Presencial
EUR 90.000 - 130.000
Hybrid work model
25 holiday days per year
Career development opportunities
+1
ai engineer for Smart TV recommendations
ai engineer for Smart TV recommendations

Enfint • Barcelona

Presencial
EUR 55.000 - 85.000
Competitive compensation
Private health insurance
International work environment
+2
Ultralytics LLM Engineer
Ultralytics LLM Engineer

Tmoose • Madrid

Presencial
EUR 85.000 - 125.000
Competitive salary
24 days paid vacation
Home setup allowance
+2