ML Platform Engineer: GPU Orchestration & Scale

Mistral

Palo Alto (CA)

On-site

USD 180,000 - 280,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
Parental leave
Retirement plans
Relocation support
Wellness programs
Meal and transportation allowances

Job summary

Mistral is hiring to build and operate the ML platform enabling large-scale training, evaluation, and batch inference. You will create infrastructure to run distributed GPU workloads across clusters and regions, with a focus on reliability and self-service for researchers and engineers.

Join a fast-moving, frontier-AI team that values ownership, observability, and scalable capacity management across heterogeneous hardware and multi-cluster environments.

Qualifications

  • 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.

Responsibilities

  • Build the ML Platform: Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.

Skills

Python or Go
Kubernetes
Distributed systems
GPU/ML workloads

Tools

Kueue
Karpenter
Volcano
Kyverno

Job description

Mistral is hiring to build and operate the ML platform enabling large-scale training, evaluation, and batch inference. You will create infrastructure to run distributed GPU workloads across clusters and regions, with a focus on reliability and self-service for researchers and engineers.

Join a fast-moving, frontier-AI team that values ownership, observability, and scalable capacity management across heterogeneous hardware and multi-cluster environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Platform Engineer — Scalable GPU & Kubernetes
ML Platform Engineer — Scalable GPU & Kubernetes

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
Research Engineer, ML Platform
Research Engineer, ML Platform

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
Research Platform Engineer
Research Platform Engineer

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
Founding ML Platforms Engineer: GPU Orchestration & Scale
Founding ML Platforms Engineer: GPU Orchestration & Scale

Cumulus Labs (YC W26) • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
Sovereign AI Compute Engineer | Linux, Kubernetes & GPUs
Sovereign AI Compute Engineer | Linux, Kubernetes & GPUs

Mistral • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
Healthcare coverage
Parental leave
Relocation support
+2
AI Infrastructure Engineer — Large-Scale Sandboxing & MLOps
AI Infrastructure Engineer — Large-Scale Sandboxing & MLOps

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 250,000
ML Infra Engineer: GPU Orchestration & Observability
ML Infra Engineer: GPU Orchestration & Observability

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000
GPU Platform Engineer — Self-Serve ML Compute
GPU Platform Engineer — Self-Serve ML Compute

B Capital • United States

On-site
USD 180,000 - 230,000
Research Engineer (ML) — Scale AI Pipelines & Production
Research Engineer (ML) — Scale AI Pipelines & Production

Mistral • San Francisco (CA)

On-site
USD 150,000 - 210,000