Member of Technical Staff, GPU / ML Systems

SkyPilot

San Mateo (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Medical/Dental/Vision coverage
Autonomy & ownership

Job summary

SkyPilot is seeking an engineer to own the GPU and ML-systems layer that powers frontier AI workloads. You will shape accelerator scheduling, utilization, and the serving path across clouds and Kubernetes, delivering a fast, cost-efficient AI stack.

You will deepen integrations with vLLM, PyTorch, Slime and related frameworks while enabling scalable training, inference, and multi-cluster serving. You’ll work with a team moving cutting-edge AI compute forward at scale.

Qualifications

  • Hands-on experience with GPU or accelerator systems and ML training/inference infrastructure.
  • Strong Python, and comfort with systems-level and GPU-adjacent details.
  • Experience operating large-scale training or high-throughput inference in production.

Responsibilities

  • Own GPU scheduling, utilization and health across clouds and Kubernetes with real-time monitoring and automatic recovery.
  • Build optimizations for training and serving: large-scale pre-training, checkpointing, and multi-cluster serving.
  • Deepen integrations with vLLM, PyTorch, CUDA and the ML frameworks used for pre-training and high-throughput inference.

Skills

GPU systems
Python
ML training/inference
Systems-level programming
Kubernetes

Tools

vLLM
PyTorch
CUDA
Slime
Kueue
KServe

Job description

About SkyPilot

SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single \"AI supercomputer.\"

SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.

The role

SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.

What you'll do
  • Own GPU scheduling, utilization and health: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.
  • Build optimizations for training and serving: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.
  • Make the AI stack run great out of the box: deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference.
What we're looking for
  • Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
  • Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
  • You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
  • Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
  • You care about squeezing most from the compute available to you
  • Experience operating large-scale training or high-throughput inference in production
What we offer
  • Competitive compensation and equity
  • Comprehensive medical, dental, vision coverage for you and your dependents
  • The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
  • A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
  • Gourmet lunch & dinner for the team to do their best work

Location: San Mateo, CA. Remote will be considered for exceptional candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Engineering Manager
Member of Technical Staff, Engineering Manager

SkyPilot • San Mateo (CA)

On-site
USD 200,000 - 320,000
Equity
Medical coverage
Autonomy and ownership
+2
MTS, Engineering Manager
MTS, Engineering Manager

Socket.dev • San Mateo (CA)

On-site
USD 210,000 - 260,000
Competitive compensation and equity
Medical, dental, vision coverage for你u
Autonomy and ownership
+2
GPU ML Systems Engineer — AI Compute Orchestration (Remote Considered)
GPU ML Systems Engineer — AI Compute Orchestration (Remote Considered)

SkyPilot • San Mateo (CA)

On-site
USD 180,000 - 280,000
Equity
Medical/Dental/Vision coverage
Autonomy & ownership
Senior, Staff Backend Engineer - Distributed System
Senior, Staff Backend Engineer - Distributed System

MetAntz • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Hybrid work model
Equity
Competitive salary
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff, Product Engineering
Member of Technical Staff, Product Engineering

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 225,000 - 275,000
Relocation assistance
Equity
Member of Technical Staff, Product Engineering
Member of Technical Staff, Product Engineering

SF Tensor • San Francisco (CA)

On-site
USD 225,000 - 275,000