Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)

NJF Global Holdings Ltd

New York (NY)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A next-generation quantitative trading firm in New York seeks a distributed-systems architect to design and operate bare-metal RDMA fabrics for over 1,000 heterogeneous accelerators. This position demands expertise in managing large-scale GPU clusters while ensuring p99.9 latency in high-volume trading environments. The ideal candidate will have a strong understanding of distributed systems architecture and experience with custom scheduling plugins. Join a dynamic firm where every second counts to enhance our trading edge.

Qualifications

  • Experience designing bare-metal RDMA fabrics for large-scale systems.
  • Expertise in custom scheduler plugins for mixed workloads.
  • Proven track record in zero-downtime upgrades and fault tolerance.

Responsibilities

  • Build and operate the cluster substrate ensuring p99.9 latency SLAs.
  • Design and optimize cost/utilization for inference cycles.
  • Manage observability stack guaranteeing pod-wide latency.

Skills

End-to-end ownership of large-scale heterogeneous GPU clusters
Distributed systems architecture expertise

Tools

RDMA
NVIDIA NVLink/NVSwitch
Slurm
Kubernetes

Job description

Are you the distributed-systems architect who has designed bare-metal RDMA fabrics for 1,000+ heterogeneous accelerators while guaranteeing p99.9 latency SLAs under live market volatility?

Core Mandate

Build and operate the physical + logical cluster substrate that makes the mission possible at scale.

What You’ll Own
  • Bare-metal / RDMA fabric design for 1,000+ GPU/TPU heterogeneous nodes (NVIDIA MIG, multi-instance GPU, TPU v5p pods).
  • Custom scheduler plugins (Slurm + Kubernetes-native) for mixed RL training/inference workloads with dynamic bin-packing.
  • Zero-downtime rolling upgrades, fault tolerance under network partitions, consensus for stateful long-sequence memory.
  • Observability stack (CUPTI + DCGM + custom eBPF) guaranteeing p99.9 latency SLAs pod-wide.
  • Cost/utilisation optimisation that directly improves P&L per inference cycle.
Must-Have Expertise
  • End-to-end ownership of large-scale heterogeneous GPU clusters (topology → orchestration → reliability at 1,000+ accelerator scale).
  • Distributed systems architecture expertise (consensus, eventual consistency, zero-downtime upgrades under extreme load).
Preferred

RDMA, NVIDIA NVLink/NVSwitch fabric tuning, custom Slurm/K8s plugins, large-scale Lustre-style long-sequence storage.

Context

A next-generation quantitative trading firm where inference latency is the only remaining alpha edge. The system ingests several million high-dimensional market‑microstructure events per second (L3 order‑book deltas, trade prints, venue feeds, order‑flow imbalance, cross‑asset signals) into stateful RL agents (actor‑critic with persistent memory) + Bayesian inference pipelines (variational GPs, Bayesian transformers, SVI nets) that maintain exact long‑range dependencies over hundreds of millions to billions of timesteps — including nanosecond‑precise microstructural signature recurrence 357 days prior at full ns timestamp resolution. The entire ingestion → feature engineering → forward pass → risk‑gates pipeline must hit p99.9 guaranteed sub‑millisecond end‑to‑end latency across heterogeneous GPU/TPU clusters with tensor/pipeline parallelism and selective FPGA/ASIC offload. Every nanosecond shaved directly increases captured P&L.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000