AI Infrastructure Engineer (GPU) - Remote EMEA

Pragmatike

Madrid

Presencial

EUR 60.000 - 80.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Work from home flexibility
Inclusive recruitment process
Opportunity to influence core engineering decisions

Descripción de la vacante

Pragmatike is looking for an AI Infrastructure Engineer to join a fast-growing startup in Madrid, fully remote. The role involves building production-grade model serving infrastructure for AI systems, focusing on efficient ML inference platforms. Candidates should have 4+ years of relevant experience, skills in model serving frameworks like vLLM and TGI, and a strong background in container orchestration. Join a dynamic team to drive innovations in AI-native cloud services.

Formación

  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems.
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent.
  • Strong background in container orchestration and operating GPU-based workloads in production.

Responsabilidades

  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton.
  • Design and implement robust deployment pipelines for ML models.
  • Develop and maintain auto-scaling systems and intelligent request routing layers.

Conocimientos

ML Ops
Platform Engineering
SRE
Model serving frameworks (vLLM, TGI, Triton)
Container orchestration
Infrastructure-as-code tools (Terraform, Helm)
Python
Distributed systems
Performance tuning

Herramientas

Terraform
Helm
Kubeflow
MLflow

Descripción del empleo

Overview

Location: Fully remote (EMEA timezone)

Start date: ASAP

Languages: Fluent English required

Industry: Cloud Computing / AI / European Deep-Tech SaaS

About The Role

Pragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company is redefining how compute is delivered by providing GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers.

We are seeking a AI Infrastructure Engineer with strong experience in production-grade model serving and infrastructure for AI systems. This is a highly technical, hands-on role focused on building scalable, reliable, and efficient ML inference platforms powering real-time AI applications.

You will be responsible for designing and operating the core infrastructure that serves machine learning models at scale. You will work closely with infrastructure, platform, and applied AI teams to ensure high availability, low latency, and cost-efficient inference systems. Strong ownership, production mindset, and experience with distributed GPU systems are essential.

Your Responsibilities
  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent
  • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health
  • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments
  • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities
  • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment
Required Qualifications
  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
  • Strong background in container orchestration and operating GPU-based workloads in production
  • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
  • Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering
  • Ability to effectively use AI coding assistants to accelerate development and debugging workflows
  • Ownership mindset with the ability to operate independently in a remote-first environment
Preferred Qualifications
  • Experience with ML platforms such as Kubeflow, MLflow, or KubeAI
  • Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems
  • Experience with cost optimization across different GPU types and inference workloads
  • Background in early-stage startups or greenfield infrastructure projects
  • Proven experience building production systems from scratch rather than maintaining legacy platforms
Why Join Us
  • Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform
  • Build foundational ML inference systems from the ground up in a high-growth, well-funded startup
  • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
  • Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems
  • Influence core engineering decisions and define best practices that will scale with the company.

Pragmatike is committed to a fair, transparent, and inclusive recruitment process. We do not discriminate based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation.

In accordance with GDPR, your personal data will be processed lawfully, fairly, and securely, and used solely for recruitment purposes, including sharing it with our client(s) for employment consideration.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote AI Infrastructure Engineer – GPU ML Ops (EMEA)
Remote AI Infrastructure Engineer – GPU ML Ops (EMEA)

Pragmatike • Madrid

Presencial
EUR 60.000 - 80.000
Infrastructure Lead
Infrastructure Lead

adm Indicia • Esplugues de Llobregat

Presencial
EUR 90.000 - 140.000
Senior MLOps Engineer (Training & Inference Optimization)
Senior MLOps Engineer (Training & Inference Optimization)

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 70.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
AI Engineer
AI Engineer

QCS Staffing • Valencia

A distancia
EUR 35.000 - 45.000
Lead Software Engineer - Working with AI
Lead Software Engineer - Working with AI

Harnham • España

Presencial
EUR 80.000 - 120.000
Competitive salary + equity/benefits package
Flexible working arrangements
Opportunity to work with cutting-edge technologies
Infrastructure Lead
Infrastructure Lead

adm Indicia • España

Presencial
EUR 70.000 - 110.000
Infrastructure Lead
Infrastructure Lead

adm Indicia • Barcelona

Presencial
EUR 70.000 - 95.000
Senior IA/ML Engineer
Senior IA/ML Engineer

Plain Concepts • España

Presencial
EUR 50.000 - 80.000
Salary determined by market
Flexible schedule (35 hours/week)
Fully remote work (optional)
+4
Senior Applied Research Engineer | Barcelona, Spain, Hybrid
Senior Applied Research Engineer | Barcelona, Spain, Hybrid

SGI • Barcelona

Híbrido
EUR 90.000 - 130.000
Equity
Relocation support
Hybrid work
+1
AI Platform Engineer
AI Platform Engineer

Peak3 (formerly ZA Tech) • Madrid

Presencial
EUR 70.000 - 100.000