AI Researcher — Inference Optimization

Featherless AI

Poland

On-site

PLN 60,000 - 80,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI research firm in Poland seeks an AI Researcher to optimize inference systems for large-scale machine learning models. The ideal candidate has a strong background in machine learning and deep learning, along with experience in optimizing inference performance. Responsibilities include designing evaluation pipelines, collaborating with engineering teams, and translating research insights into production-ready systems. Familiarity with tools like PyTorch and Triton is essential. This role offers a unique opportunity to drive innovation in AI systems.

Qualifications

  • Strong background in machine learning, deep learning, or AI systems.
  • Hands-on experience optimizing inference for large-scale models.
  • Proficiency in Python and modern ML frameworks.

Responsibilities

  • Research and develop techniques to optimize inference performance.
  • Improve latency, throughput, memory efficiency, and cost per inference.
  • Design and evaluate model-level optimizations.

Skills

Machine learning
Deep learning
Inference optimization
Python
Communication skills

Tools

PyTorch
Triton
TensorRT
ONNX Runtime

Job description

Role Overview

We are seeking an AI Researcher with deep experience in inference optimization to design, evaluate, and deploy high-performance inference systems for large-scale machine learning models. You will work at the intersection of model architecture, systems engineering, and hardware-aware optimization, improving latency, throughput, and cost efficiency across real-world production environments.

Key Responsibilities
  • Research and develop techniques to optimize inference performance for large neural networks.

  • Improve latency, throughput, memory efficiency, and cost per inference.

  • Design and evaluate model-level optimizations (quantization, pruning, KV-cache optimization, architecture-aware simplifications).

  • Implement systems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization).

  • Benchmark inference workloads across hardware accelerators.

  • Collaborate with engineering teams to deploy optimized inference pipelines.

  • Translate research insights into production-ready improvements.

Required Qualifications
  • Strong background in machine learning, deep learning, or AI systems.

  • Hands‑on experience optimizing inference for large‑scale models.

  • Proficiency in Python and modern ML frameworks (e.g., PyTorch).

  • Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime).

  • Ability to design experiments and communicate results clearly.

Preferred / Nice-to-Have Qualifications
  • Experience deploying production inference systems at scale.

  • Familiarity with distributed and multi-GPU inference.

  • Experience contributing to open‑source ML or inference frameworks.

  • Authorship or co‑authorship of peer‑reviewed research papers in machine learning, systems, or related fields.

  • Experience working close to hardware (CUDA, ROCm, profiling tools).

What Success Looks Like
  • Measurable gains in latency, throughput, and cost efficiency.

  • Optimized inference systems running reliably in production.

  • Research ideas successfully translated into deployable systems.

  • Clear benchmarks and documentation that inform product decisions.

Relevant Research Areas (Bonus)
  • Long-context inference optimization

  • Speculative decoding

  • KV-cache compression and paging

  • Efficient decoding strategies

  • Hardware-aware inference design

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer — Inference Optimization
Machine Learning Engineer — Inference Optimization

Featherless AI • Poland

Remote
PLN 255,754 - 383,631
Competitive compensation
Meaningful equity at Series A
Real ownership over performance-critical systems
AI Researcher — Training Optimization
AI Researcher — Training Optimization

Featherless AI • Poland

Remote
PLN 80,000 - 110,000
AI Researcher — AI Architecture Research
AI Researcher — AI Architecture Research

Featherless AI • Poland

Remote
PLN 298,379 - 383,631
Competitive compensation
Meaningful equity
High ownership over research direction
AI Researcher — Distillation
AI Researcher — Distillation

Featherless AI • Poland

Remote
PLN 298,379 - 383,631
Real ownership over research direction
Strong support for publishing and open research
Access to meaningful compute and production-scale problems
AI Inference Optimization Researcher
AI Inference Optimization Researcher

Featherless AI • Poland

Remote
PLN 60,000 - 80,000
Machine Learning Engineer — AI Architecture Research
Machine Learning Engineer — AI Architecture Research

Featherless AI • Poland

Remote
PLN 80,000 - 120,000
Competitive compensation
Meaningful equity
Direct influence on technical direction
Inference Stack Engineer (M/K)
Inference Stack Engineer (M/K)

Experis ManpowerGroup Sp. z o.o. • Województwo pomorskie

Hybrid
PLN 180,000 - 260,000
Medical care
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • Poland

On-site
PLN 221,000 - 384,000
Compiler Engineer
Compiler Engineer

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Machine Learning Engineer — Training Optimization
Machine Learning Engineer — Training Optimization

Featherless AI • Poland

On-site
PLN 180,000 - 260,000
Competitive compensation
Meaningful equity