Inference & Performance Engineer @ Innowise

Innowise

Poland

On-site

PLN 180,000 - 300,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Innowise is seeking a talented ML/LLM production deployment engineer in Poland to deploy and optimize models for production inference across cloud GPU and edge targets. You will work with leading inference frameworks and implement optimization techniques to ensure latency, throughput, and memory efficiency.

The role requires strong Python, ML/LLM fundamentals, and hands-on experience with at least one inference framework; knowledge of transformer architectures is essential for scalable

Qualifications

  • Strong Python and software fundamentals.
  • Hands-on production deployment of at least one ML/LLM model.
  • Experience with at least one inference/serving framework.
  • Understanding of quantization, batching, caching, and graph/kernel optimization.
  • Ability to reason about latency, throughput, memory trade-offs.
  • Solid grasp of deep learning and transformer architectures.

Responsibilities

  • Deploy and optimize ML/LLM models for production inference across cloud GPU and edge targets.
  • Work with inference frameworks and apply optimization techniques (quantization, batching, caching).
  • Build and maintain inference infrastructure with containerization, GPU scheduling (Kubernetes), autoscaling, observability, benchmarking pipelines.

Skills

Python
ML/LLM fundamentals
Model deployment
Inference frameworks
Optimization techniques
Latency optimization
Transformer concepts

Tools

vLLM
Triton
TensorRT-LLM
ONNX Runtime
llama.cpp
ggml
TGI
CUDA

Job description

Core (must-have):


  • Strong Python and solid general software engineering fundamentals

  • Hands-on experience deploying at least one ML/LLM model to production inference — cloud serving or edge, either counts

  • Working knowledge of at least one inference/serving framework (vLLM, Triton, TensorRT-LLM, ONNX Runtime, llama.cpp/ggml, TGI, or similar)

  • Practical understanding of core optimization techniques: quantization, batching, caching, graph- or kernel-level optimization

  • Comfortable reasoning about latency/throughput/memory trade-offs

  • Solid grasp of deep learning fundamentals and transformer architectures


Strong plus — this is your specialization axis, not a day-one requirement:


  • Production C++ experience, especially for edge/on-device or runtime-level work

  • CUDA / GPU kernel programming exposure

  • Direct experience with llama.cpp, ggml, TensorRT-LLM, SGLang, FlashInfer, or similar low-level inference engines

  • Kubernetes / cloud infrastructure experience for GPU workloads

  • Experience with diffusion models


[Deploy and optimize ML/LLM models for production inference across cloud GPU and, on select projects, edge/on-device targets, Work with inference/serving frameworks — vLLM, Triton Inference Server, TensorRT-LLM, ONNX Runtime, or llama.cpp/ggml, depending on the project's stack, Apply optimization techniques: quantization, pruning/distillation, operator fusion, graph/kernel-level compilation, KV-cache and batching strategies, Profile and tune runtime performance — latency, throughput, memory footprint, startup time, stability under long-running sessions, Build and maintain inference infrastructure: containerized deployment, GPU scheduling (Kubernetes), autoscaling, observability, benchmarking pipelines, On select engagements: work directly in C++ inference runtimes (e.g., llama.cpp/ggml-style engines), including custom CUDA kernel work, for edge and on-device deployment, Partner with research/ML engineers to take models from prototype to production, and with client engineering teams on integration] Requirements: Python, ML, LLM, C++, CUDA, GPU, Kubernetes

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production ML Inference Engineer
Production ML Inference Engineer

Innowise • Poland

On-site
PLN 180,000 - 300,000
Python Inference Engineer
Python Inference Engineer

Gcore • Poland

Hybrid
PLN 180,000 - 240,000
Competitive compensation
Flexible working hours and hybrid or 远
Work from anywhere in the world for up
+6
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • Poland

On-site
PLN 221,000 - 384,000
GPU Software Engineer (Graphics / ML)
GPU Software Engineer (Graphics / ML)

Luxoft • Poland

On-site
PLN 180,000 - 300,000
Private medical care
Dental care
Life insurance
+1
Senior Deep Learning Inference Engineer - GPU Optimized
Senior Deep Learning Inference Engineer - GPU Optimized

NVIDIA • Poland

On-site
PLN 221,000 - 384,000
Applied Scientist (LLM)
Applied Scientist (LLM)

SQUAD Ukraine Limited • Wrocław

Hybrid
Competitive salary packages
Guaranteed paid vacation
Private medical insurance
Compiler Engineer
Compiler Engineer

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
GPU Software Engineer (Graphics / ML)
GPU Software Engineer (Graphics / ML)

Luxoft Poland • Poland

On-site
PLN 180,000 - 260,000
Private Medical & Dental care & Life保险
MyBenefit program
Internal Mobility program
+2
Compiler Engineer - Poland
Compiler Engineer - Poland

microTECH Global Limited • Poland

Hybrid
PLN 180,000 - 280,000
Technical Lead - GPU Infrastructure
Technical Lead - GPU Infrastructure

Tether • Warszawa

Hybrid
PLN 300,000 - 520,000
Remote-friendly
Global collaboration