Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)

NJF Global Holdings Ltd

New York (NY)

On-site

USD 150,000 - 200,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A next-generation quantitative trading firm in New York seeks a distributed-systems architect to design and operate bare-metal RDMA fabrics for over 1,000 heterogeneous accelerators. This position demands expertise in managing large-scale GPU clusters while ensuring p99.9 latency in high-volume trading environments. The ideal candidate will have a strong understanding of distributed systems architecture and experience with custom scheduling plugins. Join a dynamic firm where every second counts to enhance our trading edge.

Qualifications

  • Experience designing bare-metal RDMA fabrics for large-scale systems.
  • Expertise in custom scheduler plugins for mixed workloads.
  • Proven track record in zero-downtime upgrades and fault tolerance.

Responsibilities

  • Build and operate the cluster substrate ensuring p99.9 latency SLAs.
  • Design and optimize cost/utilization for inference cycles.
  • Manage observability stack guaranteeing pod-wide latency.

Skills

End-to-end ownership of large-scale heterogeneous GPU clusters
Distributed systems architecture expertise

Tools

RDMA
NVIDIA NVLink/NVSwitch
Slurm
Kubernetes

Job description

Are you the distributed-systems architect who has designed bare-metal RDMA fabrics for 1,000+ heterogeneous accelerators while guaranteeing p99.9 latency SLAs under live market volatility?

Core Mandate

Build and operate the physical + logical cluster substrate that makes the mission possible at scale.

What You’ll Own
  • Bare-metal / RDMA fabric design for 1,000+ GPU/TPU heterogeneous nodes (NVIDIA MIG, multi-instance GPU, TPU v5p pods).
  • Custom scheduler plugins (Slurm + Kubernetes-native) for mixed RL training/inference workloads with dynamic bin-packing.
  • Zero-downtime rolling upgrades, fault tolerance under network partitions, consensus for stateful long-sequence memory.
  • Observability stack (CUPTI + DCGM + custom eBPF) guaranteeing p99.9 latency SLAs pod-wide.
  • Cost/utilisation optimisation that directly improves P&L per inference cycle.
Must-Have Expertise
  • End-to-end ownership of large-scale heterogeneous GPU clusters (topology → orchestration → reliability at 1,000+ accelerator scale).
  • Distributed systems architecture expertise (consensus, eventual consistency, zero-downtime upgrades under extreme load).
Preferred

RDMA, NVIDIA NVLink/NVSwitch fabric tuning, custom Slurm/K8s plugins, large-scale Lustre-style long-sequence storage.

Context

A next-generation quantitative trading firm where inference latency is the only remaining alpha edge. The system ingests several million high-dimensional market‑microstructure events per second (L3 order‑book deltas, trade prints, venue feeds, order‑flow imbalance, cross‑asset signals) into stateful RL agents (actor‑critic with persistent memory) + Bayesian inference pipelines (variational GPs, Bayesian transformers, SVI nets) that maintain exact long‑range dependencies over hundreds of millions to billions of timesteps — including nanosecond‑precise microstructural signature recurrence 357 days prior at full ns timestamp resolution. The entire ingestion → feature engineering → forward pass → risk‑gates pipeline must hit p99.9 guaranteed sub‑millisecond end‑to‑end latency across heterogeneous GPU/TPU clusters with tensor/pipeline parallelism and selective FPGA/ASIC offload. Every nanosecond shaved directly increases captured P&L.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

On-site
USD 200,000 - 300,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
GPU Cluster Architect
GPU Cluster Architect

Jobgether SRL • United States

Remote
USD 184,000 - 318,000
Medical, dental, vision insurance
Remote work reimbursement
RSUs may be available
+3
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Equity incentives
Lead GPU Cluster Solution Architect
Lead GPU Cluster Solution Architect

Axe Compute • Miami (FL)

On-site
USD 140,000 - 170,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000