AI Infrastructure Engineer (GPU) - Remote EMEA

Pragmatike

Turkey

On-site

TRY 2,252,252 - 4,054,054

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Ownership of critical infrastructure
Foundation building in high-growth startup
Influence core engineering decisions

Job summary

Pragmatike is seeking an AI Infrastructure Engineer to build efficient ML inference platforms for AI applications. The role requires 4+ years of experience in ML Ops or SRE, familiarity with model serving frameworks, and a solid foundation in GPU workloads. You will design production-grade infrastructure and optimize performance, working remotely across EMEA. Join a fast-growing cloud startup influencing core engineering decisions and addressing modern challenges in infrastructure development.

Qualifications

  • 4+ years in ML Ops, Platform Engineering, or similar roles.
  • Hands-on experience with model serving frameworks.
  • Strong container orchestration experience.

Responsibilities

  • Build and operate production-grade model serving infrastructure.
  • Design robust deployment pipelines for ML models.
  • Optimize GPU utilization and model storage performance.
  • Manage model registries and CI/CD pipelines.

Skills

Production-grade model serving
GPU-based workloads
MLOps tooling
Infrastructure as code
Python
Distributed systems

Tools

vLLM
TGI
Triton
Terraform
Helm

Job description

Overview

Location: Fully remote (EMEA timezone)

Start date: ASAP

Languages: Fluent English required

Industry: Cloud Computing / AI / European Deep-Tech SaaS

About The Role

Pragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company is redefining how compute is delivered by providing GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers.

We are seeking a AI Infrastructure Engineer with strong experience in production-grade model serving and infrastructure for AI systems. This is a highly technical, hands-on role focused on building scalable, reliable, and efficient ML inference platforms powering real-time AI applications.

You will be responsible for designing and operating the core infrastructure that serves machine learning models at scale. You will work closely with infrastructure, platform, and applied AI teams to ensure high availability, low latency, and cost-efficient inference systems. Strong ownership, production mindset, and experience with distributed GPU systems are essential.

Your Responsibilities
  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent
  • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health
  • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments
  • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities
  • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment
Required Qualifications
  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
  • Strong background in container orchestration and operating GPU-based workloads in production
  • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
  • Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering
  • Ability to effectively use AI coding assistants to accelerate development and debugging workflows
  • Ownership mindset with the ability to operate independently in a remote-first environment
Preferred Qualifications
  • Experience with ML platforms such as Kubeflow, MLflow, or KubeAI
  • Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems
  • Experience with cost optimization across different GPU types and inference workloads
  • Background in early-stage startups or greenfield infrastructure projects
  • Proven experience building production systems from scratch rather than maintaining legacy platforms
Why Join Us
  • Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform
  • Build foundational ML inference systems from the ground up in a high-growth, well-funded startup
  • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
  • Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems
  • Influence core engineering decisions and define best practices that will scale with the company.

Pragmatike is committed to a fair, transparent, and inclusive recruitment process. We do not discriminate based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation.

In accordance with GDPR, your personal data will be processed lawfully, fairly, and securely, and used solely for recruitment purposes, including sharing it with our client(s) for employment consideration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Infrastructure Engineer - GPU/ML Ops
Remote AI Infrastructure Engineer - GPU/ML Ops

Pragmatike • Turkey

On-site
TRY 2,252,000 - 4,055,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

The Consensus • Turkey

On-site
TRY 400,000 - 800,000
Interesting and challenging work
Learning and growth opportunities
Regular team events and offsites
AI/ML Technical Architect Lead
AI/ML Technical Architect Lead

Jobgether • Turkey

On-site
TRY 9,770,000 - 13,218,000
Medical plan options
Dental and vision insurance
401(k) match
+3
AI Platform Engineering Manager - Remote Europe
AI Platform Engineering Manager - Remote Europe

Pragmatike • Turkey

On-site
TRY 6,612,000 - 9,917,000
Remote‑first role
Lead AI Engineer with Machine Learning
Lead AI Engineer with Machine Learning

EPAM Systems • Turkey

On-site
TRY 5,747,000 - 8,621,000
Continuous upskilling
Private health insurance
Learning platforms access
+1
Senior ML Backend Engineer
Senior ML Backend Engineer

Jobgether • Turkey

On-site
TRY 400,000 - 900,000
Advanced ML projects
Geospatial analytics exposure
Cloud-native infrastructure experience
+3
Generative AI Operations Engineer (GenAI Ops)
Generative AI Operations Engineer (GenAI Ops)

EPAM Systems • Turkey

On-site
TRY 2,874,000 - 4,789,000
Continuous upskilling
Private health insurance
English courses
+3
Senior Generative AI Operations (GenAI Ops) Engineer
Senior Generative AI Operations (GenAI Ops) Engineer

EPAM Systems • Turkey

On-site
TRY 5,744,000 - 8,617,000
Continuous upskilling
Diversity of tasks
Professional development support
+3
Backend Engineer (TypeScript/React + AWS Cloud) - AI / Agentic Applications
Backend Engineer (TypeScript/React + AWS Cloud) - AI / Agentic Applications

VidRush AI Studios LLP • Turkey

On-site
TRY 1,353,000 - 3,609,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Digitopia Global Consulting Limited • Fatih

Hybrid
TRY 700,000 - 1,200,000
Stock options
Hybrid working
Flexible hours
+1