ML Platform Engineer — Scalable GPU & Kubernetes

Socket.dev

Palo Alto (CA)

On-site

USD 180,000 - 280,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
Relocation support
Retirement plans
Wellness programs
Meal & transportation allowances

Job summary

Mistral is seeking an ML Platform Engineer to build and operate the platform powering large-scale training, evaluation, and batch inference. You will develop infrastructure enabling researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions.

You will own critical systems across the ML lifecycle, from workload scheduling to production operations, improving reliability and developer experience in a fast-moving frontier AI environment.

Qualifications

  • 4+ years of experience in ML infrastructure, distributed systems, or Kubernetes platform engineering.
  • Proficient in Python or Go and productive with distributed systems.
  • Strong Kubernetes knowledge: controllers, operators, CRDs, scheduling, networking, storage, resource management.
  • Familiar with GPU workloads and related technologies (PyTorch, CUDA, NCCL).

Responsibilities

  • Build the ML Platform: services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.
  • Orchestrate GPU workloads: queueing, admission control, quotas, priorities, preemption, topology-aware placement.
  • Manage compute capacity: provisioning and allocation across clusters.
  • Enable multi-cluster execution: placement based on capacity and locality.
  • Improve researcher experience with self-service workflows for launching and debugging distributed workloads.
  • Optimize performance: GPU utilization, scheduling latency, startup time, throughput.
  • Build for reliability: observability, failure recovery, capacity planning, tooling.
  • Operate what you build: participate in on-call rotations and troubleshoot across infra.

Skills

Python
Go
Kubernetes
Distributed systems
GPU infra

Tools

Kueue
Karpenter
Volcano
Kyverno

Job description

Mistral is seeking an ML Platform Engineer to build and operate the platform powering large-scale training, evaluation, and batch inference. You will develop infrastructure enabling researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions.

You will own critical systems across the ML lifecycle, from workload scheduling to production operations, improving reliability and developer experience in a fast-moving frontier AI environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Platform Engineer: GPU Orchestration & Scale
ML Platform Engineer: GPU Orchestration & Scale

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
Research Platform Engineer
Research Platform Engineer

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
Research Engineer, ML Platform
Research Engineer, ML Platform

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
Senior ML Infra Platform Engineer — Kubernetes & GPUs
Senior ML Infra Platform Engineer — Kubernetes & GPUs

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
Sovereign AI Compute Engineer | Linux, Kubernetes & GPUs
Sovereign AI Compute Engineer | Linux, Kubernetes & GPUs

Mistral • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
Healthcare coverage
Parental leave
Relocation support
+2
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000
Research Engineer (ML) — Scale AI Pipelines & Production
Research Engineer (ML) — Scale AI Pipelines & Production

Mistral • San Francisco (CA)

On-site
USD 150,000 - 210,000
ML Platform Engineer: Build Scalable Infra for Research
ML Platform Engineer: Build Scalable Infra for Research

Anysphere • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000