Senior LLM Inference Optimization Engineer

Jobgether SRL

Lavamünd

Vor Ort

EUR 148.000 - 201.000

Vollzeit

vor 9 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Verschicke keinen 08/15-Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Competitive pay
Career growth
Flexible work
Collaborative environment
Impactful projects
Advanced AI tech
International team

Zusammenfassung

Jobgether SRL is seeking a Senior Machine Learning Engineer focused on LLM and VLM inference optimization, based in Switzerland. You will drive latency, throughput, and cost improvements across model artifacts, inference engines, and deployment architectures while collaborating with cross-functional teams to deliver measurable performance gains.

You will evaluate serving configurations, diagnose regressions, and implement advanced techniques such as quantization, distillation, and speculative

Qualifikationen

  • Strong software engineering skills in Python and PyTorch.
  • Hands-on experience deploying, operating, or optimizing LLM/VLM inference systems.
  • Experience with at least one modern inference stack (vLLM, SGLang, TensorRT-LLM, Triton, NVIDIA Dynamo, Ray Serve, KServe).
  • Strong understanding of transformer inference bottlenecks (KV cache, attention, memory bandwidth, batching).

Aufgaben

  • Own optimization initiatives for model families and inference backends.
  • Evaluate inference engines and recommend practical serving configurations.
  • Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.
  • Optimize endpoints for latency, throughput, memory efficiency, GPU utilization, and cost per token.
  • Deploy, benchmark, and extend inference engines (e.g., vLLM, TensorRT-LLM, Triton).
  • Build model-compression workflows (quantization, distillation, low-bit serving).
  • Implement advanced inference techniques (speculative decoding, KV-cache optimizations, continuous batching).
  • Develop reproducible benchmark harnesses (TTFT, tokens/sec/GPU, p95/p99 latency).
  • Collaborate with kernel/platform engineers to identify bottlenecks across stack.

Kenntnisse

Python
PyTorch
LLM Inference
Performance Tuning
Benchmarking
Distributed Systems
CUDA/Triton
Open Source

Tools

vLLM
SGLang
TensorRT-LLM
Triton Server
NVIDIA Dynamo
Ray Serve
KServe
CUDA

Jobbeschreibung

Jobgether SRL is seeking a Senior Machine Learning Engineer focused on LLM and VLM inference optimization, based in Switzerland. You will drive latency, throughput, and cost improvements across model artifacts, inference engines, and deployment architectures while collaborating with cross-functional teams to deliver measurable performance gains.

You will evaluate serving configurations, diagnose regressions, and implement advanced techniques such as quantization, distillation, and speculative

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Jobgether SRL • Lavamünd

Vor Ort
EUR 148.000 - 201.000
Competitive pay
Career growth
Flexible work
+4
Senior AI Inference Architect for Scalable LLM Deployments
Senior AI Inference Architect for Scalable LLM Deployments

NVIDIA • Lavamünd

Vor Ort
EUR 120.000 - 180.000
Competitive compensation
Comprehensive benefits package
AI Solutions Engineer — Personalization & LLM Prototyping
AI Solutions Engineer — Personalization & LLM Prototyping

hello again • Leonding

Vor Ort
EUR 60.000 - 90.000
Senior AI Infrastructure Engineer - Scale LLM Training
Senior AI Infrastructure Engineer - Scale LLM Training

Jobgether • Lavamünd

Vor Ort
EUR 90.000 - 140.000
Competitive compensation
Career growth
Flexible work
+4
Senior AI Solutions Engineer - Enterprise ML to Production
Senior AI Solutions Engineer - Enterprise ML to Production

Jobgether • Lavamünd

Vor Ort
EUR 174.000 - 304.000
Competitive base salary
Comprehensive benefits
Flexible work arrangements
+3
GenAI LLM Ops Engineer — Databricks, RAG, Remote
GenAI LLM Ops Engineer — Databricks, RAG, Remote

JobsinAustria • Wien

Vor Ort
EUR 34.000 - 41.000
Home office
Flexitime
Onboarding program
+7
AI Scientist: Building Next-Gen Engineering LLMs
AI Scientist: Building Next-Gen Engineering LLMs

Mistral • Linz

Vor Ort
EUR 90.000 - 120.000
Benefits package varies by country
AI Solutions Engineer - Personalization & LLMs
AI Solutions Engineer - Personalization & LLMs

hello again GmbH • Wien

Hybrid
EUR 60.000 - 70.000
Hybrid work
Team events
Fitness memberships
LLM & AI Automation Engineer — Low-Code & Azure AI
LLM & AI Automation Engineer — Low-Code & Azure AI

SEGULA Technologies • Graz

Vor Ort
EUR 70.000 - 110.000
LLM Software Engineer – AI Data Platforms
LLM Software Engineer – AI Data Platforms

Datavisyn GmbH • Linz

Hybrid
EUR 65.000 - 75.000
Brand new headquarters
Flexible working hours
Part-time home office
+3