Production ML Inference Engineer

Innowise

Poland

On-site

PLN 180,000 - 300,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Innowise is seeking a talented ML/LLM production deployment engineer in Poland to deploy and optimize models for production inference across cloud GPU and edge targets. You will work with leading inference frameworks and implement optimization techniques to ensure latency, throughput, and memory efficiency.

The role requires strong Python, ML/LLM fundamentals, and hands-on experience with at least one inference framework; knowledge of transformer architectures is essential for scalable

Qualifications

  • Strong Python and software fundamentals.
  • Hands-on production deployment of at least one ML/LLM model.
  • Experience with at least one inference/serving framework.
  • Understanding of quantization, batching, caching, and graph/kernel optimization.
  • Ability to reason about latency, throughput, memory trade-offs.
  • Solid grasp of deep learning and transformer architectures.

Responsibilities

  • Deploy and optimize ML/LLM models for production inference across cloud GPU and edge targets.
  • Work with inference frameworks and apply optimization techniques (quantization, batching, caching).
  • Build and maintain inference infrastructure with containerization, GPU scheduling (Kubernetes), autoscaling, observability, benchmarking pipelines.

Skills

Python
ML/LLM fundamentals
Model deployment
Inference frameworks
Optimization techniques
Latency optimization
Transformer concepts

Tools

vLLM
Triton
TensorRT-LLM
ONNX Runtime
llama.cpp
ggml
TGI
CUDA

Job description

Innowise is seeking a talented ML/LLM production deployment engineer in Poland to deploy and optimize models for production inference across cloud GPU and edge targets. You will work with leading inference frameworks and implement optimization techniques to ensure latency, throughput, and memory efficiency.

The role requires strong Python, ML/LLM fundamentals, and hands-on experience with at least one inference framework; knowledge of transformer architectures is essential for scalable

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference & Performance Engineer @ Innowise
Inference & Performance Engineer @ Innowise

Innowise • Poland

On-site
PLN 180,000 - 300,000
Production AI Engineer: Python, Cloud & LLMs
Production AI Engineer: Python, Cloud & LLMs

Sii Poland • Kraków

On-site
PLN 180,000 - 280,000
Great Place to Work
Profit sharing
Medical care
Production ML Engineer — Hybrid Warsaw: MLOps & ETL
Production ML Engineer — Hybrid Warsaw: MLOps & ETL

Enfint • Warszawa

Hybrid
PLN 180,000 - 260,000
Senior ML Engineer: Production AI & LLM Pipelines
Senior ML Engineer: Production AI & LLM Pipelines

Williams Lea • Warszawa

On-site
PLN 181,000 - 246,000
Private medical insurance
Pension contributions
Referral Scheme
Production AI Engineer - Python, Cloud & LLMs
Production AI Engineer - Python, Cloud & LLMs

Sii Poland • Łódź

On-site
PLN 180,000 - 280,000
Great Place to Work
Profit sharing
Medical care
+2
MLOps Engineer for GenAI & Production ML | Hybrid
MLOps Engineer for GenAI & Production ML | Hybrid

Capgemini • Poland

Hybrid
PLN 180,000 - 270,000
Private medical care
Life insurance
Hybrid work model
+1
Senior ML Engineer, LLMs & AWS Production
Senior ML Engineer, LLMs & AWS Production

Provectus • Kraków

On-site
PLN 190,355 - 274,957
Production AI Engineer — Python, Cloud & MLOps
Production AI Engineer — Python, Cloud & MLOps

Sii Poland • Bydgoszcz

On-site
PLN 150,000 - 210,000
Great Place to Work
Solid financial situation
Contracts with the biggest brands
+9
Production AI Engineer (ML Ops & LLMs)
Production AI Engineer (ML Ops & LLMs)

Talan Group • Warszawa

Hybrid
PLN 120,000 - 180,000
Private medical insurance
Life insurance
Lunch and transport card
+4
Senior ML Engineer: Production-Grade ML Deployment
Senior ML Engineer: Production-Grade ML Deployment

GlobalLogic • Kraków

On-site
PLN 180,000 - 260,000
Flexible opportunities
Career development
DEI matters
+2