Senior Site Reliability Engineer (GPU & ML Infrastructure)

Criteo

Grenoble

Sur place

EUR 90 000 - 120 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Hybrid work model
Health benefits
Mentorship & career development

Résumé du poste

At Criteo, the Platform Core group builds the foundational infrastructure powering our global advertising platform. We are expanding with a new GPU-focused team to scale ML training and inference workloads using Ray on Kubernetes and NVIDIA Triton Inference Server.

You will join the GPU team as a Site Reliability Engineer, designing, operating, and scaling infrastructure for high-performance model serving and real-time decisioning across services.

Qualifications

  • 5+ years of experience in backend engineering, SRE, or platform engineering focusing on distributed systems.
  • Strong experience with Kubernetes including workload scheduling and dynamic provisioning.
  • Hands-on experience with GPU-based workloads in production for ML training or inference.
  • Strong software engineering skills in C#, Python, Go or similar for reliable distributed systems.
  • Experience building or operating production-grade infrastructure with performance, scalability and reliability requirements.
  • Interest in automation, observability, and scalable systems.

Responsabilités

  • Design, operate, and scale infrastructure powering ML training and inference workloads.
  • Build and operate scalable Ray clusters on Kubernetes.
  • Improve provisioning, observability, reliability, and efficiency of ray-as-a-service environments.
  • Operate and optimize large-scale inference platforms using NVIDIA Triton.

Connaissances

Kubernetes
Distributed systems
Backend engineering
C#
Python
Go
GPU workloads
SRE practices

Outils

Ray on Kubernetes
NVIDIA Triton Inference Server

Description du poste

What You’ll Do:

At Criteo, the Platform Core group builds the foundational infrastructure powering our global advertising platform. We design and operate large-scale, resilient systems supporting real-time decision-making and data processing across thousands of services.

As we expand our distributed computing and ML infrastructure capabilities, we are building a new team focused on GPU platforms and high-performance model serving technologies.

As a Site Reliability Engineer in the GPU team , you will help design, operate, and scale the infrastructure powering machine learning training and inference workloads.

You will work on technologies such as:

Ray on Kubernetes

  • Build and operate scalable Ray clusters running on Kubernetes.

  • Develop reliable self-service distributed computing platforms for ML workloads.

  • Improve provisioning, observability, reliability, and operational efficiency of ray-as-a-service environments.

NVIDIA Triton Inference Server

  • Operate and optimize large-scale inference platforms using Triton.

  • Improve latency, throughput, scalability, and GPU utilization for deep learning inference workloads.

You will collaborate closely with ML engineers, data scientists, and infrastructure teams to deliver reliable, production-grade ML platforms accelerating innovation across Criteo.

Who You Are:
  • 5+ years of experience in backend engineering, Site Reliability Engineering, or platform engineering roles focused on distributed systems.

  • Strong experience with Kubernetes, including workload scheduling, dynamic provisioning, and custom controllers/operators.

  • Hands-on experience running or optimizing GPU-based workloads in production, ideally for ML training or inference systems.

  • Strong software engineering skills in C#, Python, Go, or similar languages, with a focus on building reliable distributed systems.

  • Experience building or operating production-grade infrastructure with strong requirements around performance, scalability, and reliability.

  • Strong interest in automation, observability, and designing systems that scale efficiently under high load.

Bonus Points

  • Experience with distributed ML frameworks such as Ray or similar systems.

  • Familiarity with inference serving stacks such as NVIDIA Triton or TensorRT.

  • Experience with GPU scheduling, resource management, or multi-tenant GPU platforms.

  • Exposure to cloud-native GPU orchestration (GKE, EKS, or on-prem Kubernetes GPU clusters).

We acknowledge that many candidates may not meet every single role requirement listed above. If your experience looks a little different from our requirements but you believe that you can still bring value to the role, we’d love to see your application!

Who We Are:

We’re Criteo, the Commerce Intelligence Platform. Criteo helps businesses turn shopper signals into commerce outcomes while delivering more relevant experiences for shoppers. We use proprietary commerce intelligence and AI decisioning to drive relevance for shoppers and performance for businesses.

At Criteo, our culture is as unique as it is diverse. From our offices across the globe or from the comfort of home, our 3,600 Criteos collaborate together to build an open, impactful, and forward-thinking environment.

We foster a workplace where everyone is valued, and employment decisions are based solely on skills, qualifications, and business needs—never on non-job-related factors or legally protected characteristics.

What We Offer:

Ways of working – Our hybrid model blends home with in-office experiences, making space for both.
Grow with us – Learning, mentorship & career development programs.
Your wellbeing matters – Health benefits, wellness perks & mental health support.
A team that cares – Diverse, inclusive, and globally connected.
Fair pay & perks – Attractive salary, with performance-based rewards and family-friendly policies, plus the potential for equity depending on role and level.

Additional benefits may vary depending on the country where you work and the nature of your employment with Criteo.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior GPU Platform SRE for ML Inference on Kubernetes
Senior GPU Platform SRE for ML Inference on Kubernetes

Criteo • Grenoble

Hybride
EUR 90 000 - 120 000
Hybrid work model
Health benefits
Mentorship & career development
Senior Software Engineer (Data Infrastructure & Reliability)
Senior Software Engineer (Data Infrastructure & Reliability)

Devops Academy • France

Hybride
EUR 60 000 - 90 000
Health benefits
Mentorship & career development programs
Wellness perks
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Devops Academy • France

Hybride
EUR 40 000 - 70 000
Health benefits
Learning and mentorship programs
Wellness perks
+2
Senior Data Scientist - Product Analytics
Senior Data Scientist - Product Analytics

Criteo • Paris

Sur place
EUR 55 000 - 75 000
Health benefits
Wellness perks
Mentorship and career development programs
+2
Software Engineer Intern - Front-End or Fullstack
Software Engineer Intern - Front-End or Fullstack

Criteo • Paris

Hybride
EUR 30 000 - 40 000
Hybrid work model
Learning and career development programs
Health benefits and wellness perks
+2
Software Engineer Intern (Front-End or Fullstack)
Software Engineer Intern (Front-End or Fullstack)

ENGINEERINGUK • Paris

Hybride
EUR 30 000 - 40 000
Health benefits
Learning and mentorship programs
Flexible hybrid work model
+2
Principal ML Engineer
Principal ML Engineer

Adikteev • Paris

Hybride
EUR 130 000 - 150 000
Competitive salary
Quarterly bonus
Longevity bonus
+6
Sales Director, Performance Media EMEA
Sales Director, Performance Media EMEA

Criteo • Paris

Hybride
EUR 120 000 - 180 000
Hybrid work model
Career development
Health benefits
+3
Engineering Program Manager (Intermediate or Senior)
Engineering Program Manager (Intermediate or Senior)

Criteo • Paris

Hybride
EUR 60 000 - 80 000
Health benefits
Learning and career development programs
Wellness perks
+2
Senior Business Analyst – IIT Monetize
Senior Business Analyst – IIT Monetize

Criteo • Paris

Hybride
EUR 85 000 - 120 000
Hybrid working model
Career development programs
Health benefits
+3