Senior GPU Platform SRE for ML Inference on Kubernetes

Criteo

Grenoble

Hybride

EUR 90 000 - 120 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Hybrid work model
Health benefits
Mentorship & career development

Résumé du poste

At Criteo, the Platform Core group builds the foundational infrastructure powering our global advertising platform. We are expanding with a new GPU-focused team to scale ML training and inference workloads using Ray on Kubernetes and NVIDIA Triton Inference Server.

You will join the GPU team as a Site Reliability Engineer, designing, operating, and scaling infrastructure for high-performance model serving and real-time decisioning across services.

Qualifications

  • 5+ years of experience in backend engineering, SRE, or platform engineering focusing on distributed systems.
  • Strong experience with Kubernetes including workload scheduling and dynamic provisioning.
  • Hands-on experience with GPU-based workloads in production for ML training or inference.
  • Strong software engineering skills in C#, Python, Go or similar for reliable distributed systems.
  • Experience building or operating production-grade infrastructure with performance, scalability and reliability requirements.
  • Interest in automation, observability, and scalable systems.

Responsabilités

  • Design, operate, and scale infrastructure powering ML training and inference workloads.
  • Build and operate scalable Ray clusters on Kubernetes.
  • Improve provisioning, observability, reliability, and efficiency of ray-as-a-service environments.
  • Operate and optimize large-scale inference platforms using NVIDIA Triton.

Connaissances

Kubernetes
Distributed systems
Backend engineering
C#
Python
Go
GPU workloads
SRE practices

Outils

Ray on Kubernetes
NVIDIA Triton Inference Server

Description du poste

At Criteo, the Platform Core group builds the foundational infrastructure powering our global advertising platform. We are expanding with a new GPU-focused team to scale ML training and inference workloads using Ray on Kubernetes and NVIDIA Triton Inference Server.

You will join the GPU team as a Site Reliability Engineer, designing, operating, and scaling infrastructure for high-performance model serving and real-time decisioning across services.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior Site Reliability Engineer (GPU & ML Infrastructure)
Senior Site Reliability Engineer (GPU & ML Infrastructure)

Criteo • Grenoble

Sur place
EUR 90 000 - 120 000
Hybrid work model
Health benefits
Mentorship & career development
Hybrid GPU Cloud SRE Manager for AI Infrastructure
Hybrid GPU Cloud SRE Manager for AI Infrastructure

Scaleway • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work up to 3 days per week
Modern offices near public transport
Healthy meals at HQ and Swile lunchカード
Senior SRE - GPU Fleet Reliability & Automation
Senior SRE - GPU Fleet Reliability & Automation

Sesterce Group • Paris

Sur place
EUR 45 000 - 70 000
Senior SRE - Scale Real-Time Data Infra
Senior SRE - Scale Real-Time Data Infra

Pigment • Paris

Sur place
EUR 70 000 - 100 000
Senior Site Reliability Engineer - GPU Fleet
Senior Site Reliability Engineer - GPU Fleet

Sesterce Group • Paris

Sur place
EUR 45 000 - 70 000
GPU Cloud SRE Engineering Lead
GPU Cloud SRE Engineering Lead

Webhosting • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work
Dining service
Swile card
GPU Cloud SRE Engineering Manager | Hybrid Role
GPU Cloud SRE Engineering Manager | Hybrid Role

Scaleway • Paris

Hybride
EUR 110 000 - 140 000
Hybrid work up to 3 days remote per wk
Office near public transport
Swile meal card
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA AI • Aillas

Sur place
EUR 120 000 - 180 000
Senior Software Engineer (Data Infrastructure & Reliability)
Senior Software Engineer (Data Infrastructure & Reliability)

Devops Academy • France

Hybride
EUR 60 000 - 90 000
Health benefits
Mentorship & career development programs
Wellness perks
+1
Systems Software Engineer, Kubernetes Scale - DGX Cloud
Systems Software Engineer, Kubernetes Scale - DGX Cloud

NVIDIA • France

Sur place
EUR 60 000 - 90 000