ML Ops Engineer (EMEA Remote)

Pragmatike

Town of Italy (NY)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Pragmatike is seeking a seasoned professional in machine learning operations to build AI-native cloud services. Your role includes creating production-grade infrastructure and managing complex ML systems. Candidates must have hands-on experience with model serving frameworks and strong Python skills, along with a proven track record in distributed systems. Join a fast-scaling startup committed to sustainability and innovation in cloud computing—fully remote role available.

Qualifications

  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar roles.
  • Hands-on experience with model serving frameworks.
  • Strong background in container orchestration and GPU-based workloads.

Responsibilities

  • Build and operate production-grade model serving infrastructure.
  • Design and implement robust deployment pipelines.
  • Manage model registries and CI/CD pipelines.

Skills

ML Ops
Python
Container orchestration
Distributed systems
Infrastructure-as-code

Tools

vLLM
TGI
Triton
Terraform
Helm

Job description

Location: Fully remote (EMEA timezone)

Start date: ASAP

Languages: Fluent English required

Industry: Cloud Computing / AI / European Deep-Tech SaaS

About The Role

Pragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company is redefining how compute is delivered by providing GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers.

Your Responsibilities
  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent
  • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health
  • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments
  • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities
  • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment
Required Qualifications
  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
  • Strong background in container orchestration and operating GPU-based workloads in production
  • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
  • Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering
  • Ability to effectively use AI coding assistants to accelerate development and debugging workflows
  • Ownership mindset with the ability to operate independently in a remote-first environment
Preferred Qualifications
  • Experience with ML platforms such as Kubeflow, MLflow, or KubeAI
  • Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems
  • Experience with cost optimization across different GPU types and inference workloads
  • Background in early-stage startups or greenfield infrastructure projects
  • Proven experience building production systems from scratch rather than maintaining legacy platforms
Why Join Us
  • Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform
  • Build foundational ML inference systems from the ground up in a high-growth, well-funded startup
  • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
  • Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems
  • Influence core engineering decisions and define best practices that will scale with the company.

Pragmatike is committed to a fair, transparent, and inclusive recruitment process. We do not discriminate based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation.

In accordance with GDPR, your personal data will be processed lawfully, fairly, and securely, and used solely for recruitment purposes, including sharing it with our client(s) for employment consideration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal ML Ops Engineer
Principal ML Ops Engineer

Pragmatike • New Jersey

Hybrid
USD 130,000 - 180,000
Competitive salary & equity options
Sign-on bonus
Health, Dental, and Vision
+1
Principal ML Ops Engineer
Principal ML Ops Engineer

Pragmatike • Cambridge (MA)

Hybrid
USD 120,000 - 160,000
Competitive salary & equity options
Sign-on bonus
Health, Dental, and Vision
+1
Principal ML Ops Engineer
Principal ML Ops Engineer

Pragmatike • San Francisco (CA)

Hybrid
USD 130,000 - 175,000
Competitive salary & equity options
Sign-on bonus
Health, Dental, and Vision
+1
Principal ML Ops Engineer
Principal ML Ops Engineer

Pragmatike • Pennsylvania

Hybrid
USD 120,000 - 180,000
Competitive salary & equity options
Sign-on bonus
Health, Dental, and Vision
+1
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
MLOps Engineer (Remote)
MLOps Engineer (Remote)

Cedent • United States

Remote
USD 55,104 - 82,656
Health Insurance
Dental Insurance
Vision Insurance
ML Ops Engineer
ML Ops Engineer

Veriipro • Town of Brookfield (WI)

On-site
USD 120,000 - 160,000
Senior MLOps Engineer
Senior MLOps Engineer

AppRecode, Inc. • Town of Middletown (NY)

On-site
USD 120,000 - 160,000
Senior AI Engineer
Senior AI Engineer

Latitude • Sacramento (CA)

Hybrid
USD 200,000 - 350,000