Machine Learning Engineer — Inference Optimization

Featherless AI

Poland

On-site

PLN 255,754 - 383,631

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Meaningful equity at Series A
Real ownership over performance-critical systems

Job summary

A pioneering AI tech company in Poland is seeking a Machine Learning Engineer to enhance model inference performance at scale. You will optimize latency and cost for large-scale ML models, collaborate with research engineers to deploy architectures, and profile inference pipelines to ensure reliability. The ideal candidate has strong experience in ML optimization, a solid grasp of deep learning concepts, and is comfortable in fast-paced startup settings. Competitive compensation and equity options are offered.

Qualifications

  • Strong experience in ML inference optimization or high-performance ML systems.
  • Solid understanding of deep learning internals like attention and memory layout.
  • Hands-on experience with PyTorch (or similar) and model deployment.
  • Familiarity with GPU performance tuning techniques.

Responsibilities

  • Optimize inference latency and cost for large-scale ML models in production.
  • Profile GPU/CPU inference pipelines for memory and process bottlenecks.
  • Collaborate with research engineers to productionize new model architectures.
  • Benchmark performance across hardware like NVIDIA and AMD GPUs.

Skills

ML inference optimization
Deep learning internals
PyTorch or similar
GPU performance tuning
Scaling inference for real users
Startup environments

Tools

CUDA
Triton
TensorRT

Job description

About the Role

We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting‑edge models into fast, reliable, and cost‑efficient systems that serve real users.

This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production‑grade performance gains.

What You’ll Do
  • Optimize inference latency, throughput, and cost for large-scale ML models in production

  • Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO)

  • Implement and tune techniques such as:

    • Quantization (fp16, bf16, int8, fp8)

    • KV‑cache optimization & reuse

    • Speculative decoding, batching, and streaming

    • Model pruning or architectural simplifications for inference

  • Collaborate with research engineers to productionize new model architectures

  • Build and maintain inference‑serving systems (e.g. Triton, custom runtimes, or bespoke stacks)

  • Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups

  • Improve system reliability, observability, and cost efficiency under real workloads

What We’re Looking For
  • Strong experience in ML inference optimization or high‑performance ML systems

  • Solid understanding of deep learning internals (attention, memory layout, compute graphs)

  • Hands‑on experience with PyTorch (or similar) and model deployment

  • Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel‑level optimizations)

  • Experience scaling inference for real users (not just research benchmarks)

  • Comfortable working in fast‑moving startup environments with ownership and ambiguity

Nice to Have
  • Experience with LLM or long‑context model inference

  • Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton)

  • Experience optimizing across different hardware vendors

  • Open‑source contributions in ML systems or inference tooling

  • Background in distributed systems or low‑latency services

Why Join Us
  • Real ownership over performance‑critical systems

  • Direct impact on product reliability and unit economics

  • Close collaboration with research, infra, and product

  • Competitive compensation + meaningful equity at Series A

  • A team that cares about engineering quality, not hype

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Researcher — Inference Optimization
AI Researcher — Inference Optimization

Featherless AI • Poland

Remote
PLN 60,000 - 80,000
Machine Learning Engineer — Training Optimization
Machine Learning Engineer — Training Optimization

Featherless AI • Poland

On-site
PLN 180,000 - 260,000
Competitive compensation
Meaningful equity
Senior Machine Learning Engineer
Senior Machine Learning Engineer

NTIATIVE IT Recruitment • Warszawa

On-site
PLN 180,000 - 300,000
Machine Learning Engineer — AI Architecture Research
Machine Learning Engineer — AI Architecture Research

Featherless AI • Poland

Remote
PLN 80,000 - 120,000
Competitive compensation
Meaningful equity
Direct influence on technical direction
Machine Learning Engineer
Machine Learning Engineer

NTIATIVE IT Recruitment • Warszawa

On-site
PLN 180,000 - 240,000
AI Researcher — Training Optimization
AI Researcher — Training Optimization

Featherless AI • Poland

Remote
PLN 80,000 - 110,000
AI Researcher — Distillation
AI Researcher — Distillation

Featherless AI • Poland

Remote
PLN 298,379 - 383,631
Real ownership over research direction
Strong support for publishing and open research
Access to meaningful compute and production-scale problems
Machine Learning Engineer — Distillation
Machine Learning Engineer — Distillation

Featherless AI • Poland

Remote
PLN 255,754 - 341,005
Competitive compensation
Meaningful equity
Remote-friendly environment
ML Inference Performance Engineer — Scale & Cost
ML Inference Performance Engineer — Scale & Cost

Featherless AI • Poland

Remote
PLN 255,754 - 383,631
Competitive compensation
Meaningful equity at Series A
Real ownership over performance-critical systems
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • Poland

On-site
PLN 221,000 - 384,000