HPC Engineer

Arlequin AI

Paris

Hybride

EUR 90 000 - 130 000

Plein temps

Il y a 3 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Une candidature sur mesure pour ce poste — un CV et une lettre de motivation personnalisés qui correspondent à l’offre.

Passez les filtres ATS

Résumé du poste

Arlequin AI is hiring an experienced HPC Engineer to design and operate a cloud-native, sovereign compute platform that powers our AI products and research. You will work on GPU/CPU clusters, orchestration with Kubernetes, and end-to-end performance optimization in a hybrid/remote/onsite Paris setup.

The role focuses on scalable infrastructure, IaC, and platform engineering to enable researchers and engineers to ship experiments efficiently while maintaining security and reliability.

Qualifications

  • 5+ years of experience in infrastructure / HPC / compute platforms.
  • Advanced proficiency with Kubernetes in production.
  • Hands-on experience with GPU workloads and batch schedulers.
  • Experience operating production services with latency and availability constraints.
  • Strong IaC experience (Terraform / OpenTofu, Terragrunt).
  • Understanding of distributed computing and large-scale batch processing.

Responsabilités

  • Design, deploy, and maintain infrastructure on Scaleway and automate with IaC (OpenTofu, Terragrunt).
  • Operate high-performance storage and I/O for training and inference.
  • Design, deploy, and operate a Kubernetes compute cluster for heterogeneous workloads; manage scheduling and resources.
  • Identify bottlenecks across compute, network, and storage; run model serving and large-scale pipelines.
  • Monitor performance, latency, and costs; implement alerts and right-size resources.
  • Collaborate with research/product to evolve the platform roadmap and provide self-service abstractions.

Connaissances

Kubernetes
GPU workloads
IaC
Terragrunt
OpenTofu
Terraform
Multi-node compute

Outils

NVIDIA GPU Operator
CUDA
DCGM
SGLang
Kueue
NCCL
MPI

Description du poste

Design and operate Arlequin's cloud-native, sovereign HPC platform powering our AI products and research

5+ years, Senior

Full-time

Paris (hybrid / full-remote / on-site flexible)

About Arlequin

Arlequin AI is both a research lab in topological deep learning and an AI platform. Our first product, HuDex, converts massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes instead of days, providing an operational advantage for governments and businesses.
~30 people, post-seed, Series A in progress. On-site in Paris.

The role

We are building a cloud-native, sovereign, and end-to-end automated high-performance computing platform to support the scientific computing workloads of our products and research. Our products are also deployed on-premise, in air-gapped environments, for clients with the strongest security requirements.

The Platform team owns the technical foundations Arlequin AI runs on: DevOps, MLOps, FinOps, FieldOps, Security/Compliance, compute and IT. Our role is not to build infrastructure for its own sake — we build self-service tools so engineers can ship without waiting on us, Forward Deployed Engineers can deploy clients without our help, C-levels understand the real cost of what they sell, and researchers don’t have to deal with engineering questions to run their experiments. Today we are a team of 3, and we aim to grow the team to 8 people by December.

As an HPC Engineer, you join the Platform team to design and operate our compute platform covering all of Arlequin’s compute workloads, ensuring a performant, scalable, and reliable platform — 100% cloud-native and orchestrated by Kubernetes. Concretely, you operate GPU and CPU clusters on Kubernetes and optimize end-to-end performance (GPU, high-throughput network interconnect, high-performance storage), while building the tools and abstractions that make teams autonomous in using, launching, and monitoring their workloads.

Your responsibilities

Infrastructure: design, deploy, and maintain infrastructure on Scaleway; industrialize IaC with OpenTofu and Terragrunt

Data storage: design and operate high-performance storage (parallel/distributed file systems, data caches, object storage) and optimize end-to-end I/O for training and inference

Compute platform: design, deploy, and operate a Kubernetes compute cluster sized for heterogeneous workloads; set up scheduling (queues, priorities, gang scheduling, fair-sharing); optimize GPU sharing and utilization; deploy and maintain device plugins

Workloads: identify and eliminate bottlenecks across compute, network, and I/O; design and operate model serving (real-time, streaming, batch, progressive rollouts); operate large-scale batch pipelines and training pipelines with the research and MLOps teams

Instrumentation & optimization: monitor compute and GPUs; instrument inference services (latency, throughput, error rate, SLOs); optimize compute costs with unit metrics; configure alerting; improve utilization and right-sizing

Platform engineering: collaborate with research/product on the platform roadmap; support capacity planning; build self-service abstractions; document the platform and runbooks

On-call: no rotation today, incidents are handled during business hours. When a rotation becomes necessary, it will never exceed one week on-call out of five, and will be compensated

Stack

GPU: NVIDIA GPU Operator, CUDA, DCGM

Serving & batch: SGLang, Kueue

Distributed computing: NCCL, MPI, RDMA / RoCE

Storage: Blob Storage

Languages: Python, Bash

What we’re looking for

5+ years of experience in infrastructure / HPC / compute platforms, including production experience

Advanced proficiency with Kubernetes in a production environment

Hands-on experience with GPU workloads and batch schedulers: scheduling, resource management, and sharing

Experience operating a production service under latency and availability constraints: autoscaling, load management, SLOs, incident management

Strong IaC experience (Terraform / OpenTofu, ideally Terragrunt)

Good understanding of distributed computing (multi-node / multi-GPU) and large-scale batch processing

Autonomy and a strong sense of ownership over the production scope, excellent communication, a feedback culture, technical curiosity, and pragmatism

Bonus

Experience with Scaleway

In-depth knowledge of the NVIDIA / ROCm / TPU ecosystems

Hands-on experience with inference optimization: quantization, continuous batching, KV-cache / prefix caching, speculative decoding, prefill/decode disaggregation

RDMA / InfiniBand / RoCE and low-latency networking

FinOps awareness, high-performance networking and storage knowledge, background in scientific computing or research, experience optimizing compute code

Process

First-fit interview (30 min)

Technical interview with the Platform team (coding + design)

Meeting with the hiring manager and the founders

Offer

Package

Competitive compensation including equity stake

Flexible remote policy: hybrid / full-remote / full on-site

Training, conference, and certification budget, with time dedicated to CNCF/LF open source contributions

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Site Reliability Engineer
Site Reliability Engineer

Arlequin AI • Paris

Hybride
EUR 90 000 - 130 000
Equity stake
Flexible remote policy
Training, conference, and cert budget
Senior HPC Platform Engineer — Cloud-Native Compute (Paris)
Senior HPC Platform Engineer — Cloud-Native Compute (Paris)

Arlequin AI • Paris

Hybride
EUR 90 000 - 130 000
Senior Product Manager
Senior Product Manager

Arlequin AI • Paris

Hybride
EUR 90 000 - 135 000
Equity stake
On-site in Paris
Datacenter Hardware Engineer, HPC
Datacenter Hardware Engineer, HPC

Mistral • Paris

Sur place
EUR 75 000 - 95 000
Health insurance
Transportation allowance
Sport allowance
+3
Infrastructure & Platform Architect
Infrastructure & Platform Architect

Business At Work • Paris

Sur place
EUR 70 000 - 110 000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alice & Bob • Paris

Hybride
EUR 85 000 - 120 000
BSPCE plan
Direct IP compensation bonuses
Flexible remote policy up to 40%
+5
Pre-Sales Solutions Architect - AI & GPU Infrastructure
Pre-Sales Solutions Architect - AI & GPU Infrastructure

Scaleway • Lille

Hybride
EUR 85 000 - 110 000
Hybrid work up to 3 days per week
International environment with diverse
Hardware Platform Operation team leader H/F
Hardware Platform Operation team leader H/F

Alice & Bob • Aubervilliers

Hybride
EUR 90 000 - 130 000
BSPCE plan
Flexible remote policy 40%
Parental plan with benefits
+2
Hardware Platform Operation team leader H/F
Hardware Platform Operation team leader H/F

Alice & Bob SAS • Aubervilliers

Hybride
EUR 110 000 - 135 000
BSPCE plan
Remote policy flexible jusqu’à 40%
Parental plan
+1
Data Center Engineer
Data Center Engineer

Thor • Paris

Sur place
EUR 70 000 - 100 000