Founding Engineer - ML Platforms

Cumulus Labs (YC W26)

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cumulus Labs (YC W26) is building the systems layer for AI infrastructure in San Francisco. We seek an ML Platforms Engineer to build and run the orchestration layer that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running efficiently.

You'll own the GPU orchestrator end-to-end, implement multi-tenant primitives, and drive observability with metrics, logs, and traces at scale.

Qualifications

  • Strong fundamentals in data structures and algorithms, with distributed systems knowledge.
  • Real production experience maintaining high-availability systems.
  • Excellent design instincts and ability to trade off decisions.
  • Fast learner capable of deep work; Go/Kubernetes experience is a plus, not required.
  • Proficiency with modern AI coding tools (Claude Code) to move fast with rigor.
  • Interested in in-person collaboration in a small team solving novel problems.

Responsibilities

  • Build and extend GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime
  • Design multi-tenant primitives: quotas, isolation, usage metering, tenant-facing inference gateway
  • Own observability: metrics, logs, traces at scale
  • Debug hard, systems-level issues from scheduling to GPU memory and networking
  • Make architectural decisions and own systems end-to-end
  • Ship fast and work directly with the founder

Skills

Data structures
Algorithms
Distributed systems
Production experience
Design instincts
Go/Kubernetes familiarity
Claude Code familiarity
In-person collaboration

Tools

Go
Kubernetes

Job description

About the role

Cumulus Labs builds the software that turns raw GPU capacity into fast, cheap, production AI. We're looking for an ML Platforms Engineer to help build and run the orchestration layer underneath our inference and agent products, the system that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running at high utilization.

What you'll do
  • Build and extend our GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime
  • Design and evolve multi-tenant primitives: quotas, isolation, usage metering, a tenant‑facing inference gateway
  • Own observability for the fleet: metrics, logs, and traces at scale
  • Debug hard, systems‑level problems across the stack, from scheduling logic down to GPU memory and networking
  • Make real architectural decisions, not just implement someone else’s design
  • Ship fast, own your systems end to end, and work directly with the founder
What we’re looking for
  • Excellent fundamentals: data structures, algorithms, distributed systems concepts, and the judgment to apply the right pattern to the right problem
  • Real production experience, ideally with systems that had to stay up and scale under load
  • Strong design instincts: you can reason about tradeoffs, not just follow a framework’s conventions
  • Fast learner who can go deep in unfamiliar territory; specific experience with Go or Kubernetes is a plus, not a requirement
  • Comfortable using modern AI coding tools (we use Claude Code heavily) to move fast without losing rigor
  • You want to work in person, in a small team, solving problems nobody has solved before
Why Cumulus

We're small, early, and building the systems layer for the next generation of AI infrastructure. You'll own real infrastructure from day one, not tickets in a backlog.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding ML Platforms Engineer: GPU Orchestration & Scale
Founding ML Platforms Engineer: GPU Orchestration & Scale

Cumulus Labs (YC W26) • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Post-training ML Engineer
Post-training ML Engineer

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000