Staff Engineer, Compute Platform & GPU Infra

Visa Hunt

New York (NY)

On-site

USD 180,000 - 240,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited vacation
Visa sponsorship
Team events

Job summary

Reflection is a research lab focused on making intelligence open and accessible. The Compute Platform team runs a Kubernetes-based platform across multiple neo-clouds, ensuring healthy compute, fault tolerance, and high availability for large GPU workloads.

You will co-design remediation strategies with training teams and drive platform evolution into multi-cloud storage and petabyte-scale data replication.

Qualifications

  • Systems-level engineering with a focus on cluster-wide behavior and maintenance.
  • Strong coding ability and a demonstrated focus on systems or GPU infrastructure.
  • Deep GPU hardware knowledge beyond standard Kubernetes, e.g., familiarity with NCCL.
  • Alignment with a K8s-first architecture.
  • Cloud storage expertise across multiple data centers, managing high-performance data products and checkpointing at scale.

Responsibilities

  • Cluster Management: Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning and rapid hardware debugging.
  • Platform Engineering: Design and iterate on our cluster management stack for workloads across large, multi-GPU fleets.
  • Monitoring & Observability: Implement comprehensive cluster-wide monitoring focusing on durability and performance benchmarking.
  • Roadmap Execution: Prepare infrastructure for next-generation GPU deployments and larger multi-cloud storage and network performance goals.

Skills

Systems-level engineering
GPU infrastructure
Kubernetes
Cloud scaling

Tools

NCCL

Job description

Reflection is a research lab focused on making intelligence open and accessible. The Compute Platform team runs a Kubernetes-based platform across multiple neo-clouds, ensuring healthy compute, fault tolerance, and high availability for large GPU workloads.

You will co-design remediation strategies with training teams and drive platform evolution into multi-cloud storage and petabyte-scale data replication.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compute Platform Lead — Multi-Cloud & GPU Scale
Compute Platform Lead — Multi-Cloud & GPU Scale

B Capital • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Staff Engineer - Large-Scale GPU Inference & RL Infra
Staff Engineer - Large-Scale GPU Inference & RL Infra

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness benefits
+5
Compute Platform Lead: Multi-Cloud, GPU Ops, & Mentorship
Compute Platform Lead: Multi-Cloud, GPU Ops, & Mentorship

Reflection AI • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 280,000
Top-tier compensation
Stock options
Health benefits
+5
Compute Platform Lead: Multi-Cloud Kubernetes & GPU
Compute Platform Lead: Multi-Cloud Kubernetes & GPU

Socket.dev • San Francisco (CA)

On-site
USD 210,000 - 320,000
Top-tier compensation and equity
Stock options
Health & wellness
+5
Member of Technical Staff - Engineering Lead, Compute Platform
Member of Technical Staff - Engineering Lead, Compute Platform

Doist • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Stock options
Health & wellness
Meals provided in office
+1
Member of Technical Staff - Compute Platform
Member of Technical Staff - Compute Platform

Visa Hunt • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Compute Platform Engineering Lead (Multi-Cloud)
Compute Platform Engineering Lead (Multi-Cloud)

Doist • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 240,000
Stock options
Health & wellness
Meals provided in office
+1
Member of Technical Staff - Engineering Lead, Compute Platform
Member of Technical Staff - Engineering Lead, Compute Platform

Reflection AI • New York (NY), San Francisco (CA)

On-site
USD 180,000 - 280,000
Top-tier compensation
Stock options
Health benefits
+5
Member of Technical Staff - Compute Platform
Member of Technical Staff - Compute Platform

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health benefits
Paid parental leave
+2