Principal Engineer, Cluster Orchestration

CoreWeave

Seattle (WA)

On-site

USD 210,000 - 320,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

CoreWeave seeks a veteran Principal Engineer to define long-term architecture for orchestration platforms, balancing performance, reliability, and cost. You will lead Kubernetes-native control planes, orchestrate workload admission and rollout, and drive multi-tenant GPU isolation strategies.

You will mentor senior engineers, influence across infrastructure and product teams, and engage with customers and open-source communities on deep technical topics as needed.

Qualifications

  • 15+ years of experience building and operating large-scale distributed systems.
  • Deep, practical knowledge of Kubernetes and Slurm internals.
  • Experience running GPU-heavy platforms for AI training, inference, or HPC workloads.
  • Strong background in Go and cloud-native systems development.

Responsibilities

  • Define the long-term architecture for CoreWeave's orchestration platforms across Kubernetes, Slurm, SUNK, Kueue, and related systems.
  • Act as a technical authority on scheduling, quota enforcement, fairness, pre-emption, and multi-tenant GPU isolation.
  • Make design decisions that balance performance, reliability, cost, and operational complexity.
  • Lead the evolution of Kubernetes-native control planes, including SUNK and custom operators.
  • Design systems that support workload admission, validation, and rollout, including model onboarding flows.

Skills

Kubernetes
Slurm
Go
Distributed systems
Leadership
Architecture

Education

Bachelor's or Master's in CS/Engineering

Tools

Kueue
Kubeflow
Argo Workflows

Job description

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.

About The Role

CoreWeave runs some of the largest GPU clusters in the world. The AI infrastructure behind those clusters determines how workloads are placed, how resources are shared, and how reliably systems perform under constant pressure.

What You'll Do
Architecture and Technical Direction
  • Define the long-term architecture for CoreWeave's orchestration platforms across Kubernetes, Slurm, SUNK, Kueue, and related systems.
  • Act as a technical authority on scheduling, quota enforcement, fairness, pre-emption, and multi-tenant GPU isolation.
  • Make design decisions that balance performance, reliability, cost, and operational complexity.
Orchestration Platform Development
  • Lead the evolution of Kubernetes-native control planes, including SUNK and custom operators.
  • Design systems that support workload admission, validation, and rollout, including model onboarding flows.
  • Identify and remove scaling limits across schedulers, control planes, registries, networking, and storage.
Reliability and Operations
  • Set standards for reliability, observability, and operational readiness across orchestration services.
  • Define SLOs, alerting, and incident response practices for platform-critical systems.
  • Ensure systems behave predictably during failures, peak load, and rapid growth.
Hands-on Engineering
  • Write and review production code for Kubernetes controllers, schedulers, admission logic, and internal tooling.
  • Measure and improve scheduling latency, container startup time, image distribution, and cold-start performance.
  • Lead architecture and design reviews across infrastructure teams.
Leadership and Influence
  • Mentor senior and staff engineers and help grow technical leaders.
  • Influence platform, infrastructure, security, and product teams through clear technical judgment.
  • Engage with customers and open-source communities on deep technical topics when needed.
Who You Are
  • 15+ years of experience building and operating large-scale distributed systems.
  • Deep, practical knowledge of Kubernetes and Slurm internals.
  • Experience running GPU-heavy platforms for AI training, inference, or HPC workloads.
  • Strong background in Go and cloud-native systems development.
  • Proven ability to set technical direction across teams without direct authority.
  • Comfortable making high-impact technical decisions in complex systems.
  • Bachelor's or Master's degree in a relevant field, or equivalent experience.
Preferred Qualifications
  • Experience with systems such as Kueue, Kubeflow, Argo Workflows, Ray, Istio, or Knative.
  • Background in ML platform engineering, model onboarding, or lifecycle management.
  • Strong understanding of scheduling strategies, pre-emption, quota enforcement, and elastic scaling.
  • Track record of operating highly reliable systems with clear SLOs and incident processes.
  • Contributions to Kubernetes, ML infrastructure, or related open-source projects.
  • Experience mentoring senior engineers and raising engineering standards.
Is This a Good Fit?

You may be a good fit if you enjoy defining long-term architecture, solving deep systems problems, and working close to the hardware layer of AI platforms. This role suits engineers who care about correctness, scale, and operational discipline, and who want their work to directly shape how AI runs in production.

Why CoreWeave?

At CoreWeave, AI infrastructure is the product. As a Principal Engineer in cluster orchestration, you will be responsible for systems that directly determine how efficiently GPUs are used, how reliably large models run, and how quickly customers can move from research to production.

This role puts you at the center of hard problems in scheduling, resource

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Cluster Orchestration
Staff Software Engineer, Cluster Orchestration

Coreweave • Bellevue (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision insurance - 100% paid
Tuition Reimbursement
401(k) with a generous employer match
+2
Senior Software Engineer, Cluster Orchestration
Senior Software Engineer, Cluster Orchestration

CoreWeave • Sunnyvale (CA)

On-site
USD 139,000 - 204,000
Medical, dental, and vision insurance
401(k) with matching
Flexible PTO
+1
Staff Software Engineer- AI Workload Orchestration
Staff Software Engineer- AI Workload Orchestration

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance—
Stock options
401(k) with employer match
+2
Staff Software Engineer, Cluster Orchestration
Staff Software Engineer, Cluster Orchestration

CoreWeave • Sunnyvale (CA)

On-site
USD 185,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Equity awards
+3
Senior Software Engineer II (IC4) – AI Workload Orchestration
Senior Software Engineer II (IC4) – AI Workload Orchestration

Coreweave • United States

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Life Insurance
ESPP
+4
Principal Engineer, Cloud Infrastructure Services
Principal Engineer, Cloud Infrastructure Services

Weights & Biases • United States

Hybrid
USD 206,000 - 303,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+3
Technical Program Manager - Cluster Orchestration & Applied Training
Technical Program Manager - Cluster Orchestration & Applied Training

CoreWeave • New York (NY)

On-site
USD 237,000 - 261,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+4
Specialist Field Engineer - Kubernetes
Specialist Field Engineer - Kubernetes

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+2
Specialist Field Engineer
Specialist Field Engineer

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with generous match
Paid parental leave
+2
Specialist Field Engineer
Specialist Field Engineer

Socket.dev • San Francisco (CA)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance
401(k) with match
Paid parental leave
+1