GPU ML Systems Engineer — AI Compute Orchestration (Remote Considered)

SkyPilot

San Mateo (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Medical/Dental/Vision coverage
Autonomy & ownership

Job summary

SkyPilot is seeking an engineer to own the GPU and ML-systems layer that powers frontier AI workloads. You will shape accelerator scheduling, utilization, and the serving path across clouds and Kubernetes, delivering a fast, cost-efficient AI stack.

You will deepen integrations with vLLM, PyTorch, Slime and related frameworks while enabling scalable training, inference, and multi-cluster serving. You’ll work with a team moving cutting-edge AI compute forward at scale.

Qualifications

  • Hands-on experience with GPU or accelerator systems and ML training/inference infrastructure.
  • Strong Python, and comfort with systems-level and GPU-adjacent details.
  • Experience operating large-scale training or high-throughput inference in production.

Responsibilities

  • Own GPU scheduling, utilization and health across clouds and Kubernetes with real-time monitoring and automatic recovery.
  • Build optimizations for training and serving: large-scale pre-training, checkpointing, and multi-cluster serving.
  • Deepen integrations with vLLM, PyTorch, CUDA and the ML frameworks used for pre-training and high-throughput inference.

Skills

GPU systems
Python
ML training/inference
Systems-level programming
Kubernetes

Tools

vLLM
PyTorch
CUDA
Slime
Kueue
KServe

Job description

SkyPilot is seeking an engineer to own the GPU and ML-systems layer that powers frontier AI workloads. You will shape accelerator scheduling, utilization, and the serving path across clouds and Kubernetes, delivering a fast, cost-efficient AI stack.

You will deepen integrations with vLLM, PyTorch, Slime and related frameworks while enabling scalable training, inference, and multi-cluster serving. You’ll work with a team moving cutting-edge AI compute forward at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, GPU / ML Systems
Member of Technical Staff, GPU / ML Systems

SkyPilot • San Mateo (CA)

On-site
USD 180,000 - 280,000
Equity
Medical/Dental/Vision coverage
Autonomy & ownership
ML Systems Engineer - Distributed AI & GPU
ML Systems Engineer - Distributed AI & GPU

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior Remote ML Infrastructure Engineer
Senior Remote ML Infrastructure Engineer

Bright Vision Technologies • Kirkland (WA)

Remote
USD 100,000 - 150,000
Engineering Manager, GPU Infrastructure & Platforms
Engineering Manager, GPU Infrastructure & Platforms

cohere • United States

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
RRSP/401K matching
+5
ML Infrastructure Engineer
ML Infrastructure Engineer

Bright Vision Technologies • Kirkland (WA)

Remote
USD 100,000 - 150,000
Applied AI/ML Engineer — Build AI at Scale (Remote)
Applied AI/ML Engineer — Build AI at Scale (Remote)

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity allocation
Health, dental, vision
Flexible PTO
+3
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Platform Engineer: Scalable GPU + Kubernetes
ML Platform Engineer: Scalable GPU + Kubernetes

Mistral • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Healthcare coverage
Parental leave
Relocation support
+2
Staff Engineer, AI Cloud Orchestration
Staff Engineer, AI Cloud Orchestration

Lambda • United States

Remote
USD 180,000 - 240,000