Compute Platform Lead: Scale Multi-Cloud GPU Clusters

Oscar Faye

New York (NY)

On-site

USD 250,000 - 360,000

Full time

24 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health cover
Parental leave
Flexible time off
Visa sponsorship
Equity

Job summary

Oscar Faye is seeking a Compute Platform Lead to own the GPU compute layer that powers frontier-scale model training. You will lead a Kubernetes-based platform spanning multiple cloud providers, tackling scheduling, health, and performance challenges at scale while staying hands-on to ship code.

You will build a small, highly capable team, shape architecture for auto-remediation and topology-aware scheduling, negotiate major compute deals, and guide the lab toward larger, petabyte-scale storage

Qualifications

  • Proven track record of building and leading infrastructure or systems teams while staying hands-on.
  • Deep systems engineering experience with how whole clusters behave and how to maintain them.
  • Strong coding ability and the credibility to earn a senior team's technical trust.
  • Depth in at least one of orchestration, storage or GPU hardware; GPU knowledge beyond standard Kubernetes (e.g. NCCL) is a plus.
  • Experience with high-performance storage across data centres and with checkpointing at scale is a plus.
  • Commercial judgement managing vendors and closing significant deals.

Responsibilities

  • Build, mentor and grow a team of strong systems engineers, scaling towards about 10.
  • Own fleet reliability and availability: scheduling across clouds, cluster management, and getting ready for next-generation GPUs and much larger clusters.
  • Make the architecture calls on automatic remediation, topology-aware scheduling, capacity planning, hardware debugging, and fleet-wide monitoring and benchmarking.
  • Work with the training teams to design fault tolerance, node health checks and remediation together.
  • Own vendor relationships, including negotiating and running major compute deals.
  • Over time, take on multi-cloud storage, petabyte-scale data replication and GPU-to-GPU network performance.

Skills

Infrastructure leadership
Systems engineering
Coding ability
Orchestration / GPU hardware
Vendor management

Job description

Oscar Faye is seeking a Compute Platform Lead to own the GPU compute layer that powers frontier-scale model training. You will lead a Kubernetes-based platform spanning multiple cloud providers, tackling scheduling, health, and performance challenges at scale while staying hands-on to ship code.

You will build a small, highly capable team, shape architecture for auto-remediation and topology-aware scheduling, negotiate major compute deals, and guide the lab toward larger, petabyte-scale storage

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Compute Platform Lead: Multi-Cloud Systems & GPU
Compute Platform Lead: Multi-Cloud Systems & GPU

reflectionai • San Francisco (CA), New York (NY)

On-site
USD 200,000 - 280,000
Top-tier compensation
Stock options
Health & wellness
+4
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

Reflection AI Ltd • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Compute Platform Lead — Multi-Cloud & GPU Scale
Compute Platform Lead — Multi-Cloud & GPU Scale

Reflection AI Ltd • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Compute Platform Lead: Multi-Cloud & GPU Systems
Compute Platform Lead: Multi-Cloud & GPU Systems

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 300,000
Top-tier compensation
Stock options
Health & wellness benefits
+2
Compute Platform Lead | Frontier AI Research Lab - New York, SF or London
Compute Platform Lead | Frontier AI Research Lab - New York, SF or London

Oscar Faye • New York (NY)

On-site
USD 250,000 - 360,000
Health cover
Parental leave
Flexible time off
+2
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • United States

Remote
USD 170,000 - 250,000
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems • United States

Remote
USD 140,000 - 230,000
Platform Engineering Lead — GPU & Kubernetes
Platform Engineering Lead — GPU & Kubernetes

Volta • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Retirement plan
+2