ML Infrastructure Engineer – GPU Compute Platform

Remanence

Paris (TX)

Hybrid

USD 124,000 - 186,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Relocation support
Hybrid work setup
Equity

Job summary

Remanence in Europe seeks a Member of Technical Staff - Infrastructure to own the compute and execution platform for training, inference and evaluation. You will ensure GPU resources are productive and research workflows are easy to operate and debug.

You will build and run GPU clusters, scheduling, networking and storage for distributed ML workloads; develop the execution platform for many concurrent environments; create data and artifact pipelines and improve observability, latency and

Qualifications

  • Strong systems fundamentals and careful reasoning about concurrency, resources and failure recovery.
  • Ability to make complex infrastructure understandable and dependable for users.
  • Prior ML infrastructure experience is valuable.

Responsibilities

  • Build and operate GPU clusters, job scheduling, networking and storage for distributed ML workloads.
  • Develop the execution platform for large numbers of concurrent environments, including sandbox isolation, resource limits, retries, and state recovery.
  • Build reliable data and artifact pipelines for datasets, trajectories, model weights, and checkpoints.
  • Own platform observability and recovery; improve capacity allocation, startup latency, and reliability through measurable changes.
  • Provide researchers with reproducible environments and simple tools to launch, inspect, and debug experiments.

Skills

Systems fundamentals
Concurrency
Distributed ML
Observability

Tools

Linux
Docker
Kubernetes
Slurm
Terraform
Python
Go
Ray
PostgreSQL
Prometheus
Grafana
OpenTelemetry
NVIDIA DCGM

Job description

Remanence in Europe seeks a Member of Technical Staff - Infrastructure to own the compute and execution platform for training, inference and evaluation. You will ensure GPU resources are productive and research workflows are easy to operate and debug.

You will build and run GPU clusters, scheduling, networking and storage for distributed ML workloads; develop the execution platform for many concurrent environments; create data and artifact pipelines and improve observability, latency and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Infrastructure
Member of Technical Staff — Infrastructure

Remanence • Paris (TX)

Hybrid
USD 124,000 - 186,000
Visa sponsorship
Relocation support
Hybrid work setup
+1
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
AI Platform Engineer: GPU Infra & MLOps
AI Platform Engineer: GPU Infra & MLOps

ICE Clear Europe Limited • Georgia

Hybrid
USD 140,000 - 200,000
Staff ML Platform Engineer: Scale GPU Pipelines
Staff ML Platform Engineer: Scale GPU Pipelines

SimplyHired • United States

Remote
USD 180,000 - 240,000
ML Infrastructure Engineer — Multi-Cluster GPU & Training
ML Infrastructure Engineer — Multi-Cluster GPU & Training

Prior Labs GmbH • North Canton (OH), New York (NY)

On-site
USD 180,000 - 260,000
Remote-friendly
Offsite team events
Global offices Berlin Freiburg NewYork
Remote GPU & ML Infrastructure Engineer
Remote GPU & ML Infrastructure Engineer

ConsultBae India Private limited • United States

Remote
USD 150,000 - 210,000
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
ML Platform Engineer: Build GPU-Scale Infra & Research
ML Platform Engineer: Build GPU-Scale Infra & Research

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
GPU Infrastructure Lead - Europe (Remote)
GPU Infrastructure Lead - Europe (Remote)

TechShack • United States

Remote
USD 180,000 - 260,000