Senior ML Ops Engineer — Remote, Scalable GPU Inference

Pragmatike

Italia

On-site

EUR 90,000 - 120,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Pragmatike is seeking a skilled ML Ops Engineer to design, build and operate scalable model serving for AI workloads in a fully remote, Europe/EMEA-based role. You will collaborate with infra, platform and AI teams to ensure low latency, high availability and cost-efficient inference across distributed GPU systems.

Responsibilities include deploying robust pipelines, auto-scaling, multi-model serving, and observability.

Qualifications

  • 4+ years in ML Ops or similar infrastructure roles.
  • Hands-on with model serving frameworks and production systems.
  • Production-grade GPU workloads experience and orchestration.
  • Python and infra-as-code tools (Terraform/Helm).
  • Strong distributed systems knowledge and reliability focus.
  • Ability to work independently in a remote-first environment.

Responsibilities

  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent.
  • Design and implement robust deployment pipelines with blue/green and canary rollouts for ML models.
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers.
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance.
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health.
  • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments.
  • Own the full lifecycle of ML systems from development through production, including operational support and on‑call responsibilities.
  • Define engineering best practices and contribute to platform scalability in a fast‑moving startup environment.

Skills

ML Ops
Model Serving
Kubernetes
Terraform
Python
CI/CD for ML
Distributed Systems
Remote Ownership
Remote Work

Tools

Kubeflow
MLflow
KubeAI

Job description

Pragmatike is seeking a skilled ML Ops Engineer to design, build and operate scalable model serving for AI workloads in a fully remote, Europe/EMEA-based role. You will collaborate with infra, platform and AI teams to ensure low latency, high availability and cost-efficient inference across distributed GPU systems.

Responsibilities include deploying robust pipelines, auto-scaling, multi-model serving, and observability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Solutions Engineer: AI & GPU Cloud (Remote)
Senior Solutions Engineer: AI & GPU Cloud (Remote)

Pragmatike • Italy

Remote
EUR 90,000 - 140,000
AI Platform Engineer: Scale GPU Jobs & Reproducible ML
AI Platform Engineer: Scale GPU Jobs & Reproducible ML

Reply • Torino

On-site
EUR 50,000 - 75,000
Staff ML Platform Engineer — Scale GPU Pipelines
Staff ML Platform Engineer — Scale GPU Pipelines

Qualcomm • Roma

Hybrid
EUR 50,000 - 80,000
ML Platform Software Engineer - Qualcomm, flexible on location anywhere in Europe
ML Platform Software Engineer - Qualcomm, flexible on location anywhere in Europe

Qualcomm • Roma

Hybrid
EUR 50,000 - 80,000
Senior ML Engineer – Generative AI & MLOps Lead
Senior ML Engineer – Generative AI & MLOps Lead

NTT DATA Europe & Latam • Roma

On-site
EUR 90,000 - 120,000
ML Engineer — Production ML & Microservices (Remote)
ML Engineer — Production ML & Microservices (Remote)

Helloprima • Milano

Hybrid
EUR 60,000 - 90,000
Private healthcare
Gym discounts
Wellbeing program
+5
Senior Solutions Engineer – Enterprise Compute
Senior Solutions Engineer – Enterprise Compute

Pragmatike • Italy

Remote
EUR 90,000 - 140,000
GenAI Platforms Engineer — Real-Time ML Systems (Remote)
GenAI Platforms Engineer — Real-Time ML Systems (Remote)

Skillvue • Milano

Remote
EUR 70,000 - 90,000
Competitive compensation
Flexible work
Budget for conferences and training
+1
Senior ML Engineer — GenAI & Production ML Lead
Senior ML Engineer — GenAI & Production ML Lead

NTT DATA Europe & Latam • Bari

On-site
EUR 90,000 - 120,000
ML Engineer - Production AI for Scalable Insurance (Remote)
ML Engineer - Production AI for Scalable Insurance (Remote)

helloprima • Milano

Hybrid
EUR 65,000 - 110,000
Private healthcare
Gym discounts
Wellbeing programs
+3