Founding Engineer — ML Platforms Engineer

cumulus labs

San Francisco (CA)

On-site

USD 140,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Cumulus Labs builds the software that turns raw GPU capacity into fast, cheap, production AI. We are hiring an ML Platforms Engineer to build and run the orchestration layer that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running at high utilization.

We care more about how you think than which languages are on your resume. Our stack today includes Go, Kubernetes, and Terraform, and we value someone who can walk into any part of a production system and make

Qualifications

  • Strong fundamentals in data structures and algorithms.
  • Experience with production systems and scaling under load.
  • Ability to reason about tradeoffs and architectural decisions.

Responsibilities

  • Build and extend GPU orchestrator for scheduling and allocation.
  • Design multi-tenant primitives including quotas and metering.
  • Own observability: metrics, logs, traces at scale.
  • Debug hard systems-level issues across stack.
  • Make architectural decisions and own end-to-end systems.
  • Ship fast and collaborate directly with the founder.

Skills

Data structures
Algorithms
Distributed systems
Production experience
Go
Kubernetes
AI coding tools
Architectural thinking

Job description

About the role

Cumulus Labs builds the software that turns raw GPU capacity into fast, cheap, production AI. We're looking for an ML Platforms Engineer to help build and run the orchestration layer underneath our inference and agent products, the system that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running at high utilization.

We care more about how you think than which languages are on your resume. Our stack today includes Go, Kubernetes, and Terraform, but we're looking for someone who can walk into any part of a production system, understand it, and make it better, not someone who only knows one toolchain.

What you'll do
  • Build and extend our GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime

  • Design and evolve multi-tenant primitives: quotas, isolation, usage metering, a tenant-facing inference gateway

  • Own observability for the fleet: metrics, logs, and traces at scale

  • Debug hard, systems-level problems across the stack, from scheduling logic down to GPU memory and networking

  • Make real architectural decisions, not just implement someone else's design

  • Ship fast, own your systems end to end, and work directly with the founder

What we're looking for
  • Excellent fundamentals: data structures, algorithms, distributed systems concepts, and the judgment to apply the right pattern to the right problem

  • Real production experience, ideally with systems that had to stay up and scale under load

  • Strong design instincts: you can reason about tradeoffs, not just follow a framework's conventions

  • Fast learner who can go deep in unfamiliar territory; specific experience with Go or Kubernetes is a plus, not a requirement

  • Comfortable using modern AI coding tools (we use Claude Code heavily) to move fast without losing rigor

  • You want to work in person, in a small team, solving problems nobody has solved before

Why Cumulus

We're small, early, and building the systems layer for the next generation of AI infrastructure. You'll own real infrastructure from day one, not tickets in a backlog.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding ML Platforms Engineer — GPU Orchestrator
Founding ML Platforms Engineer — GPU Orchestrator

cumulus labs • San Francisco (CA)

On-site
USD 140,000 - 180,000
Member of Technical Staff - Infrastructure
Member of Technical Staff - Infrastructure

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

On-site
USD 180,000 - 260,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Platform Engineer
Platform Engineer

Reactor • San Francisco (CA)

On-site
USD 190,000 - 260,000
Visa sponsorship
Relocation assistance
Health, dental, vision
+1