Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Get past ATS filters
Job summary
A next-generation quantitative trading firm in New York seeks a distributed-systems architect to design and operate bare-metal RDMA fabrics for over 1,000 heterogeneous accelerators. This position demands expertise in managing large-scale GPU clusters while ensuring p99.9 latency in high-volume trading environments. The ideal candidate will have a strong understanding of distributed systems architecture and experience with custom scheduling plugins. Join a dynamic firm where every second counts to enhance our trading edge.
Qualifications
Experience designing bare-metal RDMA fabrics for large-scale systems.
Expertise in custom scheduler plugins for mixed workloads.
Proven track record in zero-downtime upgrades and fault tolerance.
Responsibilities
Build and operate the cluster substrate ensuring p99.9 latency SLAs.
Design and optimize cost/utilization for inference cycles.
End-to-end ownership of large-scale heterogeneous GPU clusters
Distributed systems architecture expertise
Tools
RDMA
NVIDIA NVLink/NVSwitch
Slurm
Kubernetes
Job description
A next-generation quantitative trading firm in New York seeks a distributed-systems architect to design and operate bare-metal RDMA fabrics for over 1,000 heterogeneous accelerators. This position demands expertise in managing large-scale GPU clusters while ensuring p99.9 latency in high-volume trading environments. The ideal candidate will have a strong understanding of distributed systems architecture and experience with custom scheduling plugins. Join a dynamic firm where every second counts to enhance our trading edge.