Staff Platform Engineer - Multi-Cloud GPU & Kubernetes

Reflection AI Ltd

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited PTO
Visa sponsorship
Team off-sites

Job summary

Reflection AI Ltd. in the United States (New York) is seeking a Platform Engineer for our Compute Platform team. You will help manage a Kubernetes-based, multi-cloud cluster across multiple neo-clouds, focusing on health checks, remediation, and scalable GPU workloads.

Collaborate with training teams to design fault tolerance, node health checks, and performance strategies while shaping the infrastructure roadmap for future GPU deployments.

Qualifications

  • Systems engineering with a focus on cluster-wide behavior and maintenance.
  • Strong coding ability and focus on systems or GPU infrastructure.
  • Deep GPU hardware knowledge beyond standard Kubernetes, NCCL familiarity.
  • Alignment with a Kubernetes-first architecture.
  • Cloud storage expertise across multiple data centers and datasets at scale.

Responsibilities

  • Cluster management and automatic remediation tooling.
  • Platform engineering for large multi-GPU workloads.
  • Implement cluster-wide monitoring and performance benchmarks.
  • Prepare infra for next-gen GPU deployments and multi-cloud storage.

Skills

Systems engineering
GPU infrastructure
Kubernetes
Cloud storage

Job description

Reflection AI Ltd. in the United States (New York) is seeking a Platform Engineer for our Compute Platform team. You will help manage a Kubernetes-based, multi-cloud cluster across multiple neo-clouds, focusing on health checks, remediation, and scalable GPU workloads.

Collaborate with training teams to design fault tolerance, node health checks, and performance strategies while shaping the infrastructure roadmap for future GPU deployments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Compute Platform Lead: Multi-Cloud & GPU Systems
Compute Platform Lead: Multi-Cloud & GPU Systems

Reflection AI Ltd • New York (NY)

On-site
USD 180,000 - 300,000
Top-tier compensation
Stock options
Health & wellness benefits
+2
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Compute Platform Lead — Multi-Cloud & GPU Scale
Compute Platform Lead — Multi-Cloud & GPU Scale

B Capital • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)
Multi-Cloud HPC Platform Architect (Kubernetes & GPUs)

EPAM Systems Inc • United States

Remote
USD 140,000 - 230,000
Compute Platform Engineer - GPU & Multi-Cloud Infra
Compute Platform Engineer - GPU & Multi-Cloud Infra

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health benefits
Paid parental leave
+2
Senior Platform Engineer - GPU-Driven Multi-Cloud Infra
Senior Platform Engineer - GPU-Driven Multi-Cloud Infra

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU AI Platform Engineer - Kubernetes & DevOps
GPU AI Platform Engineer - Kubernetes & DevOps

AMD • San Jose (CA)

On-site
USD 140,000 - 170,000
AMD benefits
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Backend Platform Engineer – GPU Cluster Automation
Backend Platform Engineer – GPU Cluster Automation

TensorWave • Las Vegas (NV)

On-site
USD 140,000 - 210,000
Stock Options
Excellent health insurance
401(k)
+6
Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems Inc • United States

Remote
USD 140,000 - 230,000