Research Engineer, ML Platform

Mistral

Palo Alto, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
Parental leave
Relocation support
Wellness programs
Meal and transportation allowances

Job summary

Mistral in Palo Alto is seeking a backend-focused ML platform engineer to build and operate the infrastructure powering large-scale training, evaluation, and batch inference.

You will orchestrate GPU workloads, manage compute capacity across clusters, and create self-service workflows to simplify distributed workloads for researchers. This role requires deep Kubernetes knowledge and proficiency in Python or Go, plus GPU tech like PyTorch and CUDA in a fast-moving frontier AI environment.

Qualifications

  • 4+ years of experience in ML infrastructure or distributed systems.
  • Proficient in Python or Go and production-grade distributed systems.
  • Strong Kubernetes knowledge: controllers, scheduling, networking, storage, and resource management.
  • Experience with GPU workloads and technologies such as PyTorch, CUDA, NCCL.

Responsibilities

  • Build ML Platform APIs, services, and tooling for training, evaluation, fine-tuning, and batch inference.
  • Orchestrate GPU workloads with queues, admission control, quotas, priorities, preemption, and topology-aware placement.
  • Manage compute capacity across clusters and optimize heterogeneous GPU resource provisioning.
  • Enable multi-cluster execution based on capacity, data locality, hardware requirements, and priorities.
  • Improve researcher experience with self-service workflows for distributed workloads to launch, observe, debug, and reproduce.
  • Improve GPU utilization, scheduling latency, startup times, throughput, and infrastructure efficiency.
  • Build for reliability with observability, failure recovery, capacity planning, and operational tooling.
  • Operate what you build by participating in on-call rotations and troubleshooting across components.

Skills

Python
Go
Kubernetes
Distributed systems
GPU infrastructure

Tools

Kueue
Karpenter
Volcano
Kyverno
PyTorch
CUDA
NCCL

Job description

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

This role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. You will develop the infrastructure that enables researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions.

You will work across the full ML lifecycle, from workload scheduling and capacity management to platform APIs, observability, and production operations. You will take ownership of critical systems and help turn complex infrastructure into reliable, self-service capabilities.

What You Will Do
  • Build the ML Platform: Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.

  • Orchestrate GPU Workloads: Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement.

  • Manage Compute Capacity: Improve how heterogeneous GPU resources are provisioned, allocated, and utilized across clusters.

  • Enable Multi-Cluster Execution: Place workloads based on capacity, data locality, hardware requirements, and organizational priorities.

  • Improve Researcher Experience: Create self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce.

  • Optimize Performance: Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency.

  • Build for Reliability: Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.

  • Operate What You Build: Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.

What We're Looking For
  • Have 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.

  • Are proficient in Python or Go and comfortable working with production-grade distributed systems.

  • Have strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.

  • Understand technologies such as Kueue, Karpenter, Volcano, and Kyverno, and the problems they address in workload scheduling, provisioning, and policy enforcement.

  • Understand distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.

  • Are familiar with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.

  • Understand concepts such as quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.

  • Can diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware.

  • Care about developer experience and enjoy turning complex infrastructure into simple, reliable interfaces.

  • Thrive in an ambiguous, fast-moving environment shaped by frontier AI research.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, ML Platform
Research Engineer, ML Platform

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
Research Platform Engineer
Research Platform Engineer

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
Research Engineer, Machine Learning
Research Engineer, Machine Learning

Mistral • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer, Data Infrastructure
Research Engineer, Data Infrastructure

Socket.dev • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Competitive salary and equity
Medical/Dental/Vision coverage
401K with 6% matching
+6
AI Compute Engineer
AI Compute Engineer

Mistral • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Healthcare coverage
Parental leave
Retirement plans
+3
Operations Engineer, Fleet Health & Delivery
Operations Engineer, Fleet Health & Delivery

Lindus Health • Palo Alto (CA)

On-site
USD 140,000 - 200,000
Research Engineer, Machine Learning
Research Engineer, Machine Learning

Socket.dev • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Competitive salary and equity
Healthcare: Medical/Dental/Vision covered
Pension: 401K (6% matching)
+6
Reseach Engineer, Full Stack
Reseach Engineer, Full Stack

Socket.dev • Palo Alto (CA)

On-site
USD 140,000 - 180,000
Healthcare coverage
Parental leave
Retirement plans
+5
Research Engineer, Code Agents Infra
Research Engineer, Code Agents Infra

Socket.dev • Palo Alto (CA)

On-site
USD 210,000 - 260,000
Healthcare coverage
Relocation support
Wellness programs
Applied AI, Forward Deployed Machine Learning Engineer
Applied AI, Forward Deployed Machine Learning Engineer

Mistral • San Francisco (CA)

On-site
USD 140,000 - 190,000