Senior GPU Cluster Architect — Ultra-Low-Latency (1k+ Nodes)

NJF Global Holdings Ltd

New York (NY)

On-site

USD 150,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A next-generation quantitative trading firm in New York seeks a distributed-systems architect to design and operate bare-metal RDMA fabrics for over 1,000 heterogeneous accelerators. This position demands expertise in managing large-scale GPU clusters while ensuring p99.9 latency in high-volume trading environments. The ideal candidate will have a strong understanding of distributed systems architecture and experience with custom scheduling plugins. Join a dynamic firm where every second counts to enhance our trading edge.

Qualifications

  • Experience designing bare-metal RDMA fabrics for large-scale systems.
  • Expertise in custom scheduler plugins for mixed workloads.
  • Proven track record in zero-downtime upgrades and fault tolerance.

Responsibilities

  • Build and operate the cluster substrate ensuring p99.9 latency SLAs.
  • Design and optimize cost/utilization for inference cycles.
  • Manage observability stack guaranteeing pod-wide latency.

Skills

End-to-end ownership of large-scale heterogeneous GPU clusters
Distributed systems architecture expertise

Tools

RDMA
NVIDIA NVLink/NVSwitch
Slurm
Kubernetes

Job description

A next-generation quantitative trading firm in New York seeks a distributed-systems architect to design and operate bare-metal RDMA fabrics for over 1,000 heterogeneous accelerators. This position demands expertise in managing large-scale GPU clusters while ensuring p99.9 latency in high-volume trading environments. The ideal candidate will have a strong understanding of distributed systems architecture and experience with custom scheduling plugins. Join a dynamic firm where every second counts to enhance our trading edge.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)

NJF Global Holdings Ltd • New York (NY)

On-site
USD 150,000 - 200,000
GPU Systems Engineer - Low-Latency HPC & AI Clusters
GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Senior AI Infrastructure Engineer – GPU Systems
Senior AI Infrastructure Engineer – GPU Systems

Clockwork Systems, Inc. • Palo Alto (CA)

On-site
USD 130,000 - 180,000
Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cluster Architect – Solutions Lead
Senior GPU Cluster Architect – Solutions Lead

Axe Compute • Miami (FL)

On-site
USD 140,000 - 170,000
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
GPU Network Architect for AI/HPC Clusters
GPU Network Architect for AI/HPC Clusters

CyberCoders • Santa Clara (CA)

Hybrid
USD 200,000 - 250,000
Health Benefits
401k
Relocation assistance
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Network Architect for 10k+ GPU HPC Clusters
Senior Network Architect for 10k+ GPU HPC Clusters

AMD • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Hybrid work model
AMD benefits