Cluster Network Engineering Lead

Sesterce Group

San Francisco (CA)

On-site

USD 130,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sesterce Group in San Francisco is looking for a professional to lead the design and operation of AI cluster network fabrics. This role requires ownership of InfiniBand and RoCE architectures, focusing on performance and reliability for training workloads.

The ideal candidate will have a proven track record in hyperscale networking and the ability to engage with various hardware vendors. Strong leadership skills and expertise in RDMA systems are essential for success in this position.

Qualifications

  • Demonstrated experience in hyperscale networking or HPC fabrics.
  • Deep expertise in adaptive routing and queueing theory.
  • Strong knowledge of network behavior in distributed frameworks.

Responsibilities

  • Lead design and operation of Sesterce's AI cluster network fabrics.
  • Define long-range architecture for AI fabric evolution.
  • Establish observability standards for fabric telemetry.

Skills

Hyperscale networking
HPC fabrics
RDMA systems
Network telemetry
Device vendor engagement

Job description

You will lead the design and operation of Sesterce's AI cluster network fabrics — owning InfiniBand and RoCE architectures, RDMA performance, and fabric reliability for frontier-scale training and inference workloads.

What you will do
  • Define the long-range architecture for AI fabric evolution across RoCE, NVIDIA InfiniBand, Clos and spine-leaf topologies, mesh interconnects, and emerging optical fabric designs; lead migration from 100G to 400G, 800G, and 1.6T
  • Own congestion control strategy including PFC, ECN, DCQCN, queue management, path diversity, and routing policy for high-performance collective traffic
  • Drive topology design decisions that improve NCCL all-reduce performance, collective completion times, tail latency stability, and fault containment
  • Establish observability standards for fabric telemetry, queue behavior, packet loss, jitter, retry behavior, and end-to-end job impact; set cable plant strategy across fiber topology, optics qualification, and DAC/AOC standards
  • Lead vendor engagement for switches, optics, NICs, and fabric management tooling; define upgrade and migration playbooks; mentor principal and staff engineers and serve as the final escalation point for fabric architecture decisions
What we are looking for
  • Demonstrated experience in hyperscale networking, HPC fabrics, RDMA systems, or distributed systems networking at large scale (hundreds to hundreds of thousands of accelerators)
  • Deep expertise in ECMP, adaptive routing, queueing theory, network telemetry pipelines, InfiniBand, and optical systems
  • Proven track record designing and operating AI or HPC interconnects; strong working knowledge of network behavior under distributed training frameworks and collective communication libraries (NCCL, UCX)
  • Experience leading major technology transitions across multiple hardware generations and mixed-vendor environments
  • Ability to bring credibility with hardware vendors, datacenter teams, software platform leaders, and executive stakeholders
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Engineer
Network Engineer

Sesterce Group • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Cluster Networking Architect (RDMA & InfiniBand)
AI Cluster Networking Architect (RDMA & InfiniBand)

Sesterce Group • San Francisco (CA)

On-site
USD 130,000 - 160,000
Senior Technical Support Engineer, Network
Senior Technical Support Engineer, Network

NVIDIA • Germany (OH)

On-site
USD 120,000 - 180,000
Senior Technical Support Engineer, Network
Senior Technical Support Engineer, Network

NVIDIA • Town of Sweden (NY)

On-site
USD 120,000 - 160,000
Emerging Network Architect
Emerging Network Architect

NMC2 • Dallas (TX)

On-site
USD 100,000 - 140,000
Network Engineer - AI/HPC
Network Engineer - AI/HPC

Xai • Memphis (TN)

On-site
USD 180,000 - 240,000
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior Network Engineer — AI/HPC Fabric & Automation
Senior Network Engineer — AI/HPC Fabric & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Senior Network Solution Architect – AI Fabrics
Senior Network Solution Architect – AI Fabrics

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Network Architect
Network Architect

TechDigital Group • California (MO)

Hybrid
USD 120,000 - 160,000