ML Platform Engineer: Scalable GPU + Kubernetes

Mistral

Palo Alto, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
Parental leave
Relocation support
Wellness programs
Meal and transportation allowances

Job summary

Mistral in Palo Alto is seeking a backend-focused ML platform engineer to build and operate the infrastructure powering large-scale training, evaluation, and batch inference.

You will orchestrate GPU workloads, manage compute capacity across clusters, and create self-service workflows to simplify distributed workloads for researchers. This role requires deep Kubernetes knowledge and proficiency in Python or Go, plus GPU tech like PyTorch and CUDA in a fast-moving frontier AI environment.

Qualifications

  • 4+ years of experience in ML infrastructure or distributed systems.
  • Proficient in Python or Go and production-grade distributed systems.
  • Strong Kubernetes knowledge: controllers, scheduling, networking, storage, and resource management.
  • Experience with GPU workloads and technologies such as PyTorch, CUDA, NCCL.

Responsibilities

  • Build ML Platform APIs, services, and tooling for training, evaluation, fine-tuning, and batch inference.
  • Orchestrate GPU workloads with queues, admission control, quotas, priorities, preemption, and topology-aware placement.
  • Manage compute capacity across clusters and optimize heterogeneous GPU resource provisioning.
  • Enable multi-cluster execution based on capacity, data locality, hardware requirements, and priorities.
  • Improve researcher experience with self-service workflows for distributed workloads to launch, observe, debug, and reproduce.
  • Improve GPU utilization, scheduling latency, startup times, throughput, and infrastructure efficiency.
  • Build for reliability with observability, failure recovery, capacity planning, and operational tooling.
  • Operate what you build by participating in on-call rotations and troubleshooting across components.

Skills

Python
Go
Kubernetes
Distributed systems
GPU infrastructure

Tools

Kueue
Karpenter
Volcano
Kyverno
PyTorch
CUDA
NCCL

Job description

Mistral in Palo Alto is seeking a backend-focused ML platform engineer to build and operate the infrastructure powering large-scale training, evaluation, and batch inference.

You will orchestrate GPU workloads, manage compute capacity across clusters, and create self-service workflows to simplify distributed workloads for researchers. This role requires deep Kubernetes knowledge and proficiency in Python or Go, plus GPU tech like PyTorch and CUDA in a fast-moving frontier AI environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Platform Engineer — Scalable GPU & Kubernetes
ML Platform Engineer — Scalable GPU & Kubernetes

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
ML Platform Engineer: GPU Orchestration & Scale
ML Platform Engineer: GPU Orchestration & Scale

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
Research Engineer, ML Platform
Research Engineer, ML Platform

Mistral • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Healthcare coverage
Parental leave
Relocation support
+2
Research Engineer, ML Platform
Research Engineer, ML Platform

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
Research Platform Engineer
Research Platform Engineer

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
AI Infrastructure Engineer — Kubernetes, Linux & Cloud
AI Infrastructure Engineer — Kubernetes, Linux & Cloud

Lindus Health • Palo Alto (CA)

On-site
USD 140,000 - 200,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Platform Engineer: Build GPU-Scale Infra & Research
ML Platform Engineer: Build GPU-Scale Infra & Research

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Sovereign AI Compute Engineer | Linux, Kubernetes & GPUs
Sovereign AI Compute Engineer | Linux, Kubernetes & GPUs

Mistral • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
Healthcare coverage
Parental leave
Relocation support
+2