Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

Reflection AI Ltd

San Francisco (CA)

On-site

USD 210,000 - 290,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Comprehensive health/dental/vision
Meals provided in office
22 weeks parental leave
Unlimited PTO (US) / 30 days (UK)
Visa sponsorship
Team off-sites

Job summary

Reflection is seeking an experienced Compute Platform Lead in San Francisco to oversee a Kubernetes-based, multi-cloud compute layer used for large-scale AI training. You will build and mentor a high-performing team of systems engineers, guiding the architecture and fault-tolerance features while remaining hands-on when needed.

You will collaborate with training teams on remediation, plan for next-generation GPUs, and manage important vendor deals, ensuring reliability of the compute fleet

Qualifications

  • Experience building and growing a systems or infrastructure team while staying technically hands-on.
  • Deep systems-level engineering with focus on cluster-wide behavior and maintenance.
  • Strong coding ability with credibility to earn the technical trust of a strong team.
  • Depth in at least one of orchestration, storage, or GPU hardware; NCCL is a plus.
  • Alignment with a Kubernetes-first architecture.
  • Cloud storage expertise across data centers and datasets at scale.
  • Experience managing vendors and negotiating important deals.
  • Ability to guide strategy and execution across a multi-cloud, large-fleet environment.

Responsibilities

  • Build, mentor, and grow a high-performing team of systems engineers.
  • Provide front-line leadership for the compute fleet: multi-cloud scheduling, cluster management, and GPU deployments.
  • Stay hands-on to contribute technically as an individual contributor.
  • Manage day-to-day execution and prioritization in a fast-paced environment.
  • Guide architectural decisions for scalability, resilience, and monitoring.
  • Collaborate with training teams to co-design fault tolerance and remediation; manage vendor relationships.
  • Plan for next-generation GPUs and larger clusters, and long-term multi-cloud storage strategies.
  • Raise the bar for technical judgment, prioritization, and execution.

Skills

Leadership & People Mgmt
Systems Engineering
Kubernetes & Orchestration
GPU Hardware & Deployment
Multi-Cloud Architecture
Vendor Management
Strategic Planning
Communication & Mentorship

Tools

NCCL

Job description

Reflection is seeking an experienced Compute Platform Lead in San Francisco to oversee a Kubernetes-based, multi-cloud compute layer used for large-scale AI training. You will build and mentor a high-performing team of systems engineers, guiding the architecture and fault-tolerance features while remaining hands-on when needed.

You will collaborate with training teams on remediation, plan for next-generation GPUs, and manage important vendor deals, ensuring reliability of the compute fleet

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Compute Platform Lead: Multi-Cloud Systems & GPU
Compute Platform Lead: Multi-Cloud Systems & GPU

reflectionai • San Francisco (CA), New York (NY)

On-site
USD 200,000 - 280,000
Top-tier compensation
Stock options
Health & wellness
+4
Compute Platform Lead: Multi-Cloud & GPU Systems
Compute Platform Lead: Multi-Cloud & GPU Systems

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 300,000
Top-tier compensation
Stock options
Health & wellness benefits
+2
Compute Platform Lead — Multi-Cloud & GPU Scale
Compute Platform Lead — Multi-Cloud & GPU Scale

Reflection AI Ltd • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Compute Foundations Engineer — Kubernetes & GPU Infra
Compute Foundations Engineer — Kubernetes & GPU Infra

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 490,000
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems • United States

Remote
USD 140,000 - 230,000
Senior Platform Engineer - GPU-Driven Multi-Cloud Infra
Senior Platform Engineer - GPU-Driven Multi-Cloud Infra

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Compute Platform Lead | Kubernetes & Bare-Metal
Senior Compute Platform Lead | Kubernetes & Bare-Metal

Autonomai Recruitment • Chicago (IL)

On-site
USD 180,000 - 280,000
Senior Backend Engineer - GPU Cloud Platform
Senior Backend Engineer - GPU Cloud Platform

Lightning AI • San Francisco (CA), Seattle (WA), New York (NY)

Hybrid
USD 180,000 - 250,000
Comprehensive Health Coverage
Meaningful Equity
401(k) matching
+3
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • United States

Remote
USD 170,000 - 250,000