Platform Engineer: Multi-Cloud GPU Clusters & K8s

reflectionai

New York (NY)

On-site

USD 180,000 - 250,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness benefits
Meals provided in office
22 weeks parental leave
Unlimited vacation in US, 30 days inUK
Visa sponsorship

Job summary

Reflection is building a Kubernetes-first, multi-cloud compute platform to enable scalable GPU workloads and fault-tolerant operations. You will help design and maintain tooling for automatic remediation, capacity planning, and performance debugging across large GPU fleets.

The role emphasizes integrating with training teams to co-design fault tolerance and remediation strategies, while aligning with a multi-cloud storage and data replication roadmap.

Qualifications

  • Systems-level engineering experience focusing on cluster-wide behavior.
  • Strong coding ability with a focus on systems or GPU infrastructure.
  • Deep GPU hardware knowledge beyond standard Kubernetes (NCCL familiarity).
  • Alignment with a K8s-first architecture.
  • Cloud storage expertise across multi-data-center environments and data replication.

Responsibilities

  • Cluster Management: Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning and rapid hardware debugging.
  • Platform Engineering: Design and iterate on the cluster management stack for workloads across large, multi-GPU fleets.
  • Monitoring & Observability: Implement comprehensive cluster-wide monitoring, focusing on durability and performance benchmarking.
  • Roadmap Execution: Prepare the infrastructure for next-generation GPU deployments and larger clusters; own multi-cloud storage and data replication.

Skills

Systems-level engineering
Strong coding ability
GPU infrastructure
K8s-first architecture
Cloud storage expertise

Tools

Kubernetes
NCCL

Job description

Reflection is building a Kubernetes-first, multi-cloud compute platform to enable scalable GPU workloads and fault-tolerant operations. You will help design and maintain tooling for automatic remediation, capacity planning, and performance debugging across large GPU fleets.

The role emphasizes integrating with training teams to co-design fault tolerance and remediation strategies, while aligning with a multi-cloud storage and data replication roadmap.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Compute Platform Lead: Multi-Cloud Systems & GPU
Compute Platform Lead: Multi-Cloud Systems & GPU

reflectionai • San Francisco (CA), New York (NY)

On-site
USD 200,000 - 280,000
Top-tier compensation
Stock options
Health & wellness
+4
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Compute Platform Lead — Multi-Cloud & GPU Scale
Compute Platform Lead — Multi-Cloud & GPU Scale

B Capital • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Compute Platform Engineer - GPU & Multi-Cloud Infra
Compute Platform Engineer - GPU & Multi-Cloud Infra

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health benefits
Paid parental leave
+2
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Backend Platform Engineer – GPU Cluster Automation
Backend Platform Engineer – GPU Cluster Automation

TensorWave • Las Vegas (NV)

On-site
USD 140,000 - 210,000
Stock Options
Excellent health insurance
401(k)
+6
Backend Engineer - Kubernetes & GPU Cluster Automation
Backend Engineer - Kubernetes & GPU Cluster Automation

TensorWave Inc. • United States

Remote
USD 180,000 - 240,000
Stock Options
Medical, Dental, Vision insurance for
Company Health Savings Account
+7
Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Staff Software Engineer - GPU Fleet & Cloud Infra
Staff Software Engineer - GPU Fleet & Cloud Infra

Cloudjobs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Restricted Stock Units
Health insurance
Vision insurance
+15
Senior GPU Platform Architect (GPUaaS)
Senior GPU Platform Architect (GPUaaS)

Unify Technologies Ltd • Plano (TX)

On-site
USD 180,000 - 240,000