Get more replies from employers
Send a job-specific resume in minutes.
Are you the distributed-systems architect who has designed bare-metal RDMA fabrics for 1,000+ heterogeneous accelerators while guaranteeing p99.9 latency SLAs under live market volatility?
Build and operate the physical + logical cluster substrate that makes the mission possible at scale.
RDMA, NVIDIA NVLink/NVSwitch fabric tuning, custom Slurm/K8s plugins, large-scale Lustre-style long-sequence storage.
A next-generation quantitative trading firm where inference latency is the only remaining alpha edge. The system ingests several million high-dimensional market‑microstructure events per second (L3 order‑book deltas, trade prints, venue feeds, order‑flow imbalance, cross‑asset signals) into stateful RL agents (actor‑critic with persistent memory) + Bayesian inference pipelines (variational GPs, Bayesian transformers, SVI nets) that maintain exact long‑range dependencies over hundreds of millions to billions of timesteps — including nanosecond‑precise microstructural signature recurrence 357 days prior at full ns timestamp resolution. The entire ingestion → feature engineering → forward pass → risk‑gates pipeline must hit p99.9 guaranteed sub‑millisecond end‑to‑end latency across heterogeneous GPU/TPU clusters with tensor/pipeline parallelism and selective FPGA/ASIC offload. Every nanosecond shaved directly increases captured P&L.