Platform & AI Clusters Lead — Scale & SRE

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 450,000 - 550,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A pioneering technology company in San Francisco is seeking a Head of Platform/AI Cluster Management to lead AI and platform initiatives. Responsibilities include overseeing scheduler management and optimizing performance across cross-functional teams. The ideal candidate should have extensive experience in cluster management, especially with Slurm and Kubernetes. This role offers a competitive salary of $500,000 gross per year.

Qualifications

  • Deep expertise in cluster management and scheduling.
  • Hands-on experience with orchestration platforms.
  • Strong understanding of performance tuning and workload management.

Responsibilities

  • Oversee the scheduler/runtime layer including multi-tenancy and GPU management.
  • Lead cluster operations and incident response.
  • Deliver platform services for reliable workload execution.
  • Collaborate with infra and SRE teams to optimize efficiency.

Skills

Cluster management
Scheduling
Runtime environments
Slurm
Kubernetes
Ray
NCCL performance tuning
Workload isolation
Congestion management

Job description

A pioneering technology company in San Francisco is seeking a Head of Platform/AI Cluster Management to lead AI and platform initiatives. Responsibilities include overseeing scheduler management and optimizing performance across cross-functional teams. The ideal candidate should have extensive experience in cluster management, especially with Slurm and Kubernetes. This role offers a competitive salary of $500,000 gross per year.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Platform/AI Cluster Management - System Integrator
Head of Platform/AI Cluster Management - System Integrator

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 450,000 - 550,000
Senior AI GPU Infra SRE - Scale, Automation & Equity
Senior AI GPU Infra SRE - Scale, Automation & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Staff Platform Engineer — Scale & Reliability
Staff Platform Engineer — Scale & Reliability

Rogo • New York (NY)

On-site
USD 120,000 - 160,000
Exceptional traction with top financial institutions
Working with a world-class team
Opportunities for rapid learning and career growth
Senior AI Platform Engineer: Scale, Reliability & Leadership
Senior AI Platform Engineer: Scale, Reliability & Leadership

AlphaSense, Inc. • San Francisco (CA)

On-site
USD 178,000 - 267,000
Equity options
Generous benefits program
Head of Platform Infra & Foundations (Distributed Systems)
Head of Platform Infra & Foundations (Distributed Systems)

Cerebras • San Francisco (CA)

On-site
USD 260,000 - 380,000
AI Platform DevOps & SRE Lead
AI Platform DevOps & SRE Lead

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Platform SRE Lead – AI Spacetech, Equity & Ownership
Platform SRE Lead – AI Spacetech, Equity & Ownership

Attis • Houston (TX)

On-site
USD 130,000 - 150,000
Staff Platform Engineer - Kubernetes, Reliability, Scale
Staff Platform Engineer - Kubernetes, Reliability, Scale

Rogo • New York (NY)

Hybrid
USD 150,000 - 200,000
Senior Platform Engineer - Scalable AI Infrastructure
Senior Platform Engineer - Scalable AI Infrastructure

Anthropic • San Francisco (CA)

On-site
USD 405,000 - 485,000
Lead Platform Engineer, Kubernetes for AI Science
Lead Platform Engineer, Kubernetes for AI Science

European Recruitment BV • United States

On-site
USD 150,000 - 200,000