Founding ML Platforms Engineer: GPU Orchestration & Scale

Cumulus Labs (YC W26)

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cumulus Labs (YC W26) is building the systems layer for AI infrastructure in San Francisco. We seek an ML Platforms Engineer to build and run the orchestration layer that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running efficiently.

You'll own the GPU orchestrator end-to-end, implement multi-tenant primitives, and drive observability with metrics, logs, and traces at scale.

Qualifications

  • Strong fundamentals in data structures and algorithms, with distributed systems knowledge.
  • Real production experience maintaining high-availability systems.
  • Excellent design instincts and ability to trade off decisions.
  • Fast learner capable of deep work; Go/Kubernetes experience is a plus, not required.
  • Proficiency with modern AI coding tools (Claude Code) to move fast with rigor.
  • Interested in in-person collaboration in a small team solving novel problems.

Responsibilities

  • Build and extend GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime
  • Design multi-tenant primitives: quotas, isolation, usage metering, tenant-facing inference gateway
  • Own observability: metrics, logs, traces at scale
  • Debug hard, systems-level issues from scheduling to GPU memory and networking
  • Make architectural decisions and own systems end-to-end
  • Ship fast and work directly with the founder

Skills

Data structures
Algorithms
Distributed systems
Production experience
Design instincts
Go/Kubernetes familiarity
Claude Code familiarity
In-person collaboration

Tools

Go
Kubernetes

Job description

Cumulus Labs (YC W26) is building the systems layer for AI infrastructure in San Francisco. We seek an ML Platforms Engineer to build and run the orchestration layer that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running efficiently.

You'll own the GPU orchestrator end-to-end, implement multi-tenant primitives, and drive observability with metrics, logs, and traces at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Engineer - ML Platforms
Founding Engineer - ML Platforms

Cumulus Labs (YC W26) • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Infra Engineer: GPU Orchestration & Observability
ML Infra Engineer: GPU Orchestration & Observability

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer — GPU, Kubernetes Infra
Senior ML Platform Engineer — GPU, Kubernetes Infra

WorkGenius Group • Los Angeles (CA)

On-site
USD 117,000 - 186,000
GPU Platform Engineer — Self-Serve ML Compute
GPU Platform Engineer — Self-Serve ML Compute

B Capital • United States

On-site
USD 180,000 - 230,000
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Platform ML Infra Engineer: GPU, Kubernetes & MLOps
Platform ML Infra Engineer: GPU, Kubernetes & MLOps

Oracle • United States

On-site
USD 92,000 - 210,000
Health insurance
401(k) match
Paid time off
+2
ML Infra Engineer — Scale GPU ML Platform & Equity
ML Infra Engineer — Scale GPU ML Platform & Equity

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Founding HPC Engineer - GPU Cloud Infra & AI Orchestration
Founding HPC Engineer - GPU Cloud Infra & AI Orchestration

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
Post-training ML Engineer
Post-training ML Engineer

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000