Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital

San Francisco (CA)

On-site

USD 210,000 - 290,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier compensation
Stock options
Comprehensive health/dental/vision
Meals provided in office
22 weeks parental leave
Unlimited PTO (US) / 30 days (UK)
Visa sponsorship
Team off-sites

Job summary

Reflection is seeking an experienced Compute Platform Lead in San Francisco to oversee a Kubernetes-based, multi-cloud compute layer used for large-scale AI training. You will build and mentor a high-performing team of systems engineers, guiding the architecture and fault-tolerance features while remaining hands-on when needed.

You will collaborate with training teams on remediation, plan for next-generation GPUs, and manage important vendor deals, ensuring reliability of the compute fleet

Qualifications

  • Experience building and growing a systems or infrastructure team while staying technically hands-on.
  • Deep systems-level engineering with focus on cluster-wide behavior and maintenance.
  • Strong coding ability with credibility to earn the technical trust of a strong team.
  • Depth in at least one of orchestration, storage, or GPU hardware; NCCL is a plus.
  • Alignment with a Kubernetes-first architecture.
  • Cloud storage expertise across data centers and datasets at scale.
  • Experience managing vendors and negotiating important deals.
  • Ability to guide strategy and execution across a multi-cloud, large-fleet environment.

Responsibilities

  • Build, mentor, and grow a high-performing team of systems engineers.
  • Provide front-line leadership for the compute fleet: multi-cloud scheduling, cluster management, and GPU deployments.
  • Stay hands-on to contribute technically as an individual contributor.
  • Manage day-to-day execution and prioritization in a fast-paced environment.
  • Guide architectural decisions for scalability, resilience, and monitoring.
  • Collaborate with training teams to co-design fault tolerance and remediation; manage vendor relationships.
  • Plan for next-generation GPUs and larger clusters, and long-term multi-cloud storage strategies.
  • Raise the bar for technical judgment, prioritization, and execution.

Skills

Leadership & People Mgmt
Systems Engineering
Kubernetes & Orchestration
GPU Hardware & Deployment
Multi-Cloud Architecture
Vendor Management
Strategic Planning
Communication & Mentorship

Tools

NCCL

Job description

Reflection is seeking an experienced Compute Platform Lead in San Francisco to oversee a Kubernetes-based, multi-cloud compute layer used for large-scale AI training. You will build and mentor a high-performing team of systems engineers, guiding the architecture and fault-tolerance features while remaining hands-on when needed.

You will collaborate with training teams on remediation, plan for next-generation GPUs, and manage important vendor deals, ensuring reliability of the compute fleet

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compute Platform Lead: Multi-Cloud Kubernetes & GPU
Compute Platform Lead: Multi-Cloud Kubernetes & GPU

Socket.dev • San Francisco (CA)

On-site
USD 210,000 - 320,000
Top-tier compensation and equity
Stock options
Health & wellness
+5
Compute Platform Lead: Multi-Cloud, GPU Ops, & Mentorship
Compute Platform Lead: Multi-Cloud, GPU Ops, & Mentorship

Reflection AI • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 280,000
Top-tier compensation
Stock options
Health benefits
+5
Compute Platform Lead — Multi-Cloud & GPU Scale
Compute Platform Lead — Multi-Cloud & GPU Scale

B Capital • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Compute Platform Engineering Lead (Multi-Cloud)
Compute Platform Engineering Lead (Multi-Cloud)

Doist • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Stock options
Health & wellness
Meals provided in office
+1
Compute Platform Engineer - GPU & Multi-Cloud Infra
Compute Platform Engineer - GPU & Multi-Cloud Infra

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health benefits
Paid parental leave
+2
Staff Engineer, Compute Platform & GPU Infra
Staff Engineer, Compute Platform & GPU Infra

Visa Hunt • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Senior Platform Engineer - GPU-Driven Multi-Cloud Infra
Senior Platform Engineer - GPU-Driven Multi-Cloud Infra

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
GenAI Platform Engineer: Multi-GPU & Kubernetes Lead
GenAI Platform Engineer: Multi-GPU & Kubernetes Lead

Quantiphi • United States

On-site
USD 140,000 - 210,000
Senior Backend Engineer, GPU Cloud Infra & Kubernetes
Senior Backend Engineer, GPU Cloud Infra & Kubernetes

Socket.dev • New York (NY)

Hybrid
USD 180,000 - 250,000
Health insurance
Equity
401(k) matching
+6