Compute Platform Lead: Multi-Cloud Kubernetes & GPU

Socket.dev

San Francisco (CA)

On-site

USD 210,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier compensation and equity
Stock options
Health & wellness
Meals provided
Paid parental leave
Unlimited PTO (US)
Sponsorship support
Team off-sites

Job summary

Reflection is seeking a Compute Platform Lead to lead a Kubernetes-based compute layer spanning multi-cloud environments. You will mentor a team of systems engineers, guide architectural decisions, and stay hands-on to contribute as an IC on critical path work.

You will manage day-to-day execution, collaborate with training teams, and drive next-generation GPU deployments and larger clusters. The role demands strategic thinking and strong cross-team communication across a fast-paced research

Qualifications

  • Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on.
  • Deep systems-level engineering with focus on cluster-wide behavior and maintenance.
  • Strong coding ability and the credibility to earn the technical trust of a strong team.
  • Depth in at least one of orchestration, storage, or GPU hardware — with the ability to learn the rest.
  • Alignment with a Kubernetes-first architecture.
  • Cloud storage expertise — managing high-performance data products across multiple data centers is a plus.
  • Experience managing vendors, including negotiating and operating important deals.
  • Ability to guide strategy and drive execution across a multi-cloud, large-fleet environment.

Responsibilities

  • Build, mentor, and grow a high-performing team of systems engineers. Coach and support your reports in understanding, and pursuing, their professional growth.
  • Provide front-line leadership of engineering efforts to keep the compute fleet reliable and highly available — multi-cloud scheduling, cluster management, and the path to next-generation GPUs and increasingly larger cluster sizes.
  • Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual try.
  • Manage day-to-day execution: prioritize the team's work and manage projects in a highly dynamic, fast-paced environment.
  • Guide technical and architectural decisions, emphasizing scalability, robustness, and reliability — automatic remediation, topology-aware scheduling, capacity planning, rapid hardware debugging, and cluster-wide monitoring and performance benchmarking.
  • Work closely with our training teams to co-design fault tolerance, node health checks, and remediation, and manage the vendor relationships and important deals the compute fleet depends on.
  • Prepare the fleet for what's next: next-generation GPUs and larger clusters, and — longer term — multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.
  • Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment.

Skills

Leadership
Systems engineering
Kubernetes
GPU hardware
Multi-cloud
Vendor management
Communication
Architectural decisions
NCCL

Tools

NCCL

Job description

Reflection is seeking a Compute Platform Lead to lead a Kubernetes-based compute layer spanning multi-cloud environments. You will mentor a team of systems engineers, guide architectural decisions, and stay hands-on to contribute as an IC on critical path work.

You will manage day-to-day execution, collaborate with training teams, and drive next-generation GPU deployments and larger clusters. The role demands strategic thinking and strong cross-team communication across a fast-paced research

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compute Platform Lead: Multi-Cloud, GPU Ops, & Mentorship
Compute Platform Lead: Multi-Cloud, GPU Ops, & Mentorship

Reflection AI • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 280,000
Top-tier compensation
Stock options
Health benefits
+5
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Compute Platform Lead — Multi-Cloud & GPU Scale
Compute Platform Lead — Multi-Cloud & GPU Scale

B Capital • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Compute Platform Engineering Lead (Multi-Cloud)
Compute Platform Engineering Lead (Multi-Cloud)

Doist • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Stock options
Health & wellness
Meals provided in office
+1
Compute Platform Engineer - GPU & Multi-Cloud Infra
Compute Platform Engineer - GPU & Multi-Cloud Infra

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health benefits
Paid parental leave
+2
Staff Engineer, Compute Platform & GPU Infra
Staff Engineer, Compute Platform & GPU Infra

Visa Hunt • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Platform Engineering Lead — GPU & Kubernetes
Platform Engineering Lead — GPU & Kubernetes

Volta • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Retirement plan
+2
Platform Lead, Multi-Cloud GPU Inference & Training
Platform Lead, Multi-Cloud GPU Inference & Training

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000
GPUaaS Kubernetes Platform Engineer
GPUaaS Kubernetes Platform Engineer

Veriipro • Irving (TX)

On-site
USD 140,000 - 180,000