Turn this role into an interview — a resume and cover letter built around what this employer wants.
Innowise is seeking a talented ML/LLM production deployment engineer in Poland to deploy and optimize models for production inference across cloud GPU and edge targets. You will work with leading inference frameworks and implement optimization techniques to ensure latency, throughput, and memory efficiency.
The role requires strong Python, ML/LLM fundamentals, and hands-on experience with at least one inference framework; knowledge of transformer architectures is essential for scalable
[Deploy and optimize ML/LLM models for production inference across cloud GPU and, on select projects, edge/on-device targets, Work with inference/serving frameworks — vLLM, Triton Inference Server, TensorRT-LLM, ONNX Runtime, or llama.cpp/ggml, depending on the project's stack, Apply optimization techniques: quantization, pruning/distillation, operator fusion, graph/kernel-level compilation, KV-cache and batching strategies, Profile and tune runtime performance — latency, throughput, memory footprint, startup time, stability under long-running sessions, Build and maintain inference infrastructure: containerized deployment, GPU scheduling (Kubernetes), autoscaling, observability, benchmarking pipelines, On select engagements: work directly in C++ inference runtimes (e.g., llama.cpp/ggml-style engines), including custom CUDA kernel work, for edge and on-device deployment, Partner with research/ML engineers to take models from prototype to production, and with client engineering teams on integration] Requirements: Python, ML, LLM, C++, CUDA, GPU, Kubernetes