Member of Technical Staff - Compute Platform

Visa Hunt

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals provided
Parental leave
Unlimited vacation
Visa sponsorship
Team events

Job summary

Reflection is a research lab focused on making intelligence open and accessible. The Compute Platform team runs a Kubernetes-based platform across multiple neo-clouds, ensuring healthy compute, fault tolerance, and high availability for large GPU workloads.

You will co-design remediation strategies with training teams and drive platform evolution into multi-cloud storage and petabyte-scale data replication.

Qualifications

  • Systems-level engineering with a focus on cluster-wide behavior and maintenance.
  • Strong coding ability and a demonstrated focus on systems or GPU infrastructure.
  • Deep GPU hardware knowledge beyond standard Kubernetes, e.g., familiarity with NCCL.
  • Alignment with a K8s-first architecture.
  • Cloud storage expertise across multiple data centers, managing high-performance data products and checkpointing at scale.

Responsibilities

  • Cluster Management: Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning and rapid hardware debugging.
  • Platform Engineering: Design and iterate on our cluster management stack for workloads across large, multi-GPU fleets.
  • Monitoring & Observability: Implement comprehensive cluster-wide monitoring focusing on durability and performance benchmarking.
  • Roadmap Execution: Prepare infrastructure for next-generation GPU deployments and larger multi-cloud storage and network performance goals.

Skills

Systems-level engineering
GPU infrastructure
Kubernetes
Cloud scaling

Tools

NCCL

Job description

Our Mission

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

About the Role

Reflection's Compute Platform team specializes in keeping our compute layer healthy and highly available. We run a K8s-based platform distributed across multiple neo-clouds. We manage multi-cloud scheduling, node health, and performance debugging at this scale presents genuinely hard systems problems. More broadly, you will work closely with Reflection's training teams to co-design fault tolerance, node health checks, and remediation strategies.

What You'll Do
  • Cluster Management: Build and maintain tools for the automatic remediation, topology-aware scheduling, capacity planning and rapid hardware debugging.

  • Platform Engineering: Design and iterate on our cluster management stack for workloads across large, multi-GPU fleets

  • Monitoring & Observability: Implement comprehensive cluster-wide monitoring, focusing on durability and active performance benchmarking.

  • Roadmap Execution: Prepare the infrastructure for next-generation GPU deployments and increasingly larger cluster sizes. Long-term, you will help own multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.

What We're Looking For

Systems-level engineering experience with a focus on cluster-wide behavior and maintenance.

Strong coding ability and a demonstrated focus on systems or GPU infrastructure.

Deep GPU hardware knowledge beyond standard Kubernetes,e.g., familiarity with NCCL.

Alignment with a K8s-first architecture

Cloud storage expertise, specifically managing high-performance data products (like VAST) across multiple data centers, connecting those storage environments together and handling datasets and checkpointing at scale.

What We Offer:

We believe that to make intelligence open and accessible to all, you need to start at the foundation. Joining Reflection means building from the ground up as part of a talent-dense team. You will help define our future as a company, and help define the future of open foundational models.

We want you to do the most impactful work of your career with the confidence that you and the people you care about most are supported.

  • Top-tier compensation: Salary and equity structured to recognize and retain our talent globally.

  • Stock options: Everyone who joins and contributes to Reflection's success gets to share in the upside through stock options.

  • Health & wellness: Comprehensive medical, dental, vision, and life, with an annual wellness allowance.

  • Meals: Lunch and dinner are provided in the office daily.

  • Life & family: 22 weeks paid parental leave for all new birthing and non-birthing parents, including adoptive and surrogate journeys.

  • Vacation days: Unlimited paid time off in the U.S. and 30 days in the U.K.

  • Sponsorship support: We sponsor visas to help exceptional talent join our team and support long-term immigration pathways where applicable.

  • Team building: We have regular off-sites, happy hours, and team celebrations.

Export Control Notice: This position may require access to technology or source code subject to the U.S. Export Administration Regulations. Any offer of employment for this role may be conditioned on the Company's ability to provide the candidate with access to such technology or source code in compliance with applicable U.S. export control laws, which may require the Company to seek government authorization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Engineering Lead, Compute Platform
Member of Technical Staff - Engineering Lead, Compute Platform

B Capital • San Francisco (CA)

On-site
USD 240,000 - 360,000
Top-tier compensation
Stock options
Health & wellness
Member of Technical Staff - Compute Platform Lead
Member of Technical Staff - Compute Platform Lead

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Member of Technical Staff - Distributed Systems Engineer
Member of Technical Staff - Distributed Systems Engineer

Visa Hunt • New York (NY)

On-site
USD 140,000 - 210,000
Top-tier compensation
Stock options
Health & wellness
+5
Member of Technical Staff - Compute Platform
Member of Technical Staff - Compute Platform

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive health benefits
Paid parental leave
+2
Member of Technical Staff - Platform Foundations
Member of Technical Staff - Platform Foundations

Visa Hunt • New York (NY)

On-site
USD 150,000 - 210,000
Top-tier compensation
Stock options
Health & wellness benefits
+1
Member of Technical Staff - Engineering Lead, Compute Platform
Member of Technical Staff - Engineering Lead, Compute Platform

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 350,000
Top-tier compensation
Stock options
Health & wellness benefits
+3
Member of Technical Staff - Engineering Lead, Data Platform
Member of Technical Staff - Engineering Lead, Data Platform

Sierra Ventures • San Francisco (CA)

On-site
USD 260,000 - 380,000
Stock options
Health & wellness benefits
Meals provided in the office
+1
Member of Technical Staff - Data Platform
Member of Technical Staff - Data Platform

B Capital • San Francisco (CA)

On-site
USD 180,000 - 260,000
Top-tier compensation
Stock options
Health & wellness
+5
Member of Technical Staff - Engineering Lead, Data Platform
Member of Technical Staff - Engineering Lead, Data Platform

Reflection • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Member of Technical Staff - Engineering Lead, Data Platform
Member of Technical Staff - Engineering Lead, Data Platform

Reflection • San Francisco (CA)

On-site
USD 180,000 - 240,000
Stock options
Health & wellness
Meals in office
+2