ML Ops Engineer (EMEA Remote)

Pragmatike

Lisboa

Presencial

EUR 50 000 - 70 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Pragmatike is looking for an experienced ML Ops Engineer to build and operate production-grade model serving infrastructure for a fast-scaling cloud startup in Lisbon. You will design deployment pipelines, optimize GPU utilization, and manage the full lifecycle of ML systems. Ideal candidates should have 4+ years in ML Ops, hands-on experience with model serving frameworks, and a strong understanding of distributed systems. Join us for a fair and inclusive recruitment process and influence core engineering decisions.

Qualificações

  • 4+ years in ML Ops or similar roles focused on ML systems.
  • Hands-on experience with model serving frameworks.
  • Strong background in operating GPU-based workloads.

Responsabilidades

  • Build and operate production-grade model serving infrastructure.
  • Design robust deployment pipelines for ML models.
  • Optimize GPU utilization and memory efficiency.

Conhecimentos

ML Ops
Platform Engineering
Container orchestration
Python
Distributed systems

Ferramentas

Terraform
Triton
Kubeflow

Descrição da oferta de emprego

Location: Fully remote (EMEA timezone)

Start date: ASAP

Languages: Fluent English required

Industry: Cloud Computing / AI / European Deep-Tech SaaS

About The Role

Pragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company is redefining how compute is delivered by providing GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers.

Your Responsibilities
  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent
  • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health
  • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments
  • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities
  • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment
Required Qualifications
  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
  • Strong background in container orchestration and operating GPU-based workloads in production
  • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
  • Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering
  • Ability to effectively use AI coding assistants to accelerate development and debugging workflows
  • Ownership mindset with the ability to operate independently in a remote-first environment
Preferred Qualifications
  • Experience with ML platforms such as Kubeflow, MLflow, or KubeAI
  • Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems
  • Experience with cost optimization across different GPU types and inference workloads
  • Background in early-stage startups or greenfield infrastructure projects
  • Proven experience building production systems from scratch rather than maintaining legacy platforms
Why Join Us
  • Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform
  • Build foundational ML inference systems from the ground up in a high-growth, well-funded startup
  • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
  • Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems
  • Influence core engineering decisions and define best practices that will scale with the company.

Pragmatike is committed to a fair, transparent, and inclusive recruitment process. We do not discriminate based on age, disability, gender, gender identity or expression, marital or civil partner status, pregnancy or maternity, race, religion or belief, sex, or sexual orientation.

In accordance with GDPR, your personal data will be processed lawfully, fairly, and securely, and used solely for recruitment purposes, including sharing it with our client(s) for employment consideration.

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Lead ML Ops Engineer - Remote GPU Inference Platform
Lead ML Ops Engineer - Remote GPU Inference Platform

Pragmatike • Lisboa

Presencial
EUR 90 000 - 130 000
Remote ML Ops Engineer: GPU Cloud Infrastructure
Remote ML Ops Engineer: GPU Cloud Infrastructure

Pragmatike • Lisboa

Presencial
EUR 50 000 - 70 000
Senior MLOps
Senior MLOps

Emerging Travel Group • Portugal

Teletrabalho
EUR 60 000 - 80 000
Flexible work schedule
Remote or hybrid work options
Development and training programs
+3
AI Operations Engineer
AI Operations Engineer

OpenSpring • Lisboa

Presencial
EUR 50 000 - 70 000
Senior Cloud Platform Engineer (m/f/d)
Senior Cloud Platform Engineer (m/f/d)

Employer Brand Anchor • Porto

Híbrido
EUR 45 000 - 70 000
Competitive Salary
Remote Work
Relocation Support
+4
(Senior) ML Ops Engineer
(Senior) ML Ops Engineer

Continental • Lousado

Híbrido
EUR 40 000 - 60 000
Challenging international work environment
Flexible working model
Continuous personal development opportunities
Distributed Cloud | Machine Learning Engineer
Distributed Cloud | Machine Learning Engineer

Devoteam • Lisboa

Presencial
EUR 45 000 - 65 000
Senior ML Ops / LLM Ops Engineer
Senior ML Ops / LLM Ops Engineer

Caixa Mágica Software • Lisboa

Presencial
EUR 45 000 - 75 000
Health and Life Insurance
Social events and team buildings
Tech equipment and smartphone
+1
Distributed Cloud | AI/ML Engineer
Distributed Cloud | AI/ML Engineer

Devoteam • Lisboa

Presencial
EUR 50 000 - 70 000
Machine Learning Ops Engineer @JobCloud
Machine Learning Ops Engineer @JobCloud

TX Services • Porto

Híbrido
EUR 60 000 - 90 000
Hybrid work option
25 days annual leave
Meal allowance
+2