Compute Platform Lead — Multi-Cloud & GPU Scale

Reflection AI Ltd

San Francisco (CA)

On-site

USD 240,000 - 360,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness

Job summary

Reflection's Compute Platform team leads a Kubernetes-based, multi-cloud compute layer across neo-clouds. As Compute Platform Lead you will mentor a team of systems engineers, guide architectural decisions, and work with training teams on fault tolerance and remediation.

You will stay hands-on to contribute and ensure the compute fleet supports large-scale training runs. This role also manages vendor relationships, drives capacity planning, and prepares the fleet for next-gen GPUs and

Qualifications

  • Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on.
  • Deep systems-level engineering with a focus on cluster-wide behavior and maintenance.
  • Strong coding ability and credibility to earn technical trust of a team.
  • Depth in orchestration, storage, or GPU hardware with ability to learn the rest.

Responsibilities

  • Build, mentor, and grow a high-performing team of systems engineers.
  • Provide front-line leadership to keep the compute fleet reliable and highly available.
  • Stay hands-on to contribute as an individual in the team's stack.
  • Manage day-to-day execution and project priorities in a fast-paced environment.
  • Guide architectural decisions emphasizing scalability, robustness, and reliability.

Skills

Team leadership
Kubernetes
Multi-cloud
GPU deployment
Systems engineering
Vendor management
Architectural decisions
Coding ability

Tools

NCCL
CUDA
Prometheus

Job description

Reflection's Compute Platform team leads a Kubernetes-based, multi-cloud compute layer across neo-clouds. As Compute Platform Lead you will mentor a team of systems engineers, guide architectural decisions, and work with training teams on fault tolerance and remediation.

You will stay hands-on to contribute and ensure the compute fleet supports large-scale training runs. This role also manages vendor relationships, drives capacity planning, and prepares the fleet for next-gen GPUs and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Compute Platform Lead: Multi-Cloud & GPU Systems
Compute Platform Lead: Multi-Cloud & GPU Systems

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 300,000
Top-tier compensation
Stock options
Health & wellness benefits
+2
Compute Platform Lead: Multi-Cloud Systems & GPU
Compute Platform Lead: Multi-Cloud Systems & GPU

reflectionai • San Francisco (CA), New York (NY)

On-site
USD 200,000 - 280,000
Top-tier compensation
Stock options
Health & wellness
+4
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

Reflection AI Ltd • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes
Staff Platform Engineer - Multi-Cloud GPU & Kubernetes

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Senior Compute Platform Lead | Kubernetes & Bare-Metal
Senior Compute Platform Lead | Kubernetes & Bare-Metal

Autonomai Recruitment • Chicago (IL)

On-site
USD 180,000 - 280,000
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems • United States

Remote
USD 140,000 - 230,000
Senior Compute Platform Engineer - Multi-Cloud & Linux
Senior Compute Platform Engineer - Multi-Cloud & Linux

Uber Technologies Inc. • Seattle (WA)

On-site
USD 202,000 - 224,000
Bonus program
Equity award
401(k) plan
+1
Platform Engineering Lead — GPU & Kubernetes
Platform Engineering Lead — GPU & Kubernetes

Volta • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Retirement plan
+2
Senior Compute Platform Engineer - Multi-Cloud Scale
Senior Compute Platform Engineer - Multi-Cloud Scale

Uber • Seattle (WA)

On-site
USD 202,000 - 224,000
Bonus program
Equity award
401(k) plan
+1
Global GPU Capacity Lead (Kubernetes & Multi-Cloud)
Global GPU Capacity Lead (Kubernetes & Multi-Cloud)

Baseten • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive equity
100% medical/dental/vision coverage
Flexible PTO incl. Winter Break