Network Engineer

Sesterce Group

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sesterce Group is seeking a skilled networking engineer to design, deploy, and operate the network infrastructure for their GPU AI factories. The role includes working with InfiniBand and high-speed Ethernet fabrics and managing automation pipelines.

The ideal candidate will have a minimum of 4 years of datacenter networking experience, including familiarity with RDMA and strong Linux networking skills. Responsibilities also include troubleshooting network performance during live AI training runs.

Qualifications

  • Minimum of 4 years of datacenter networking experience.
  • Familiarity with RDMA, RoCE v2, and GPU training cluster communication.
  • Experience using network-as-code tooling.

Responsibilities

  • Design and deploy InfiniBand and high-speed Ethernet fabrics.
  • Configure and operate network equipment and manage overlays.
  • Tune transport for collective communication workloads.
  • Maintain network automation pipelines across all sites.
  • Troubleshoot performance regressions during live AI training.

Skills

Datacenter networking experience
InfiniBand or 400G/800G Ethernet knowledge
Linux networking internals
Network automation tools
Low-level diagnostics interpretation

Tools

Ansible
Terraform
Netbox

Job description

You will design, deploy, and operate the network infrastructure underpinning Sesterce's GPU AI factories across Europe — owning the full stack from physical cabling to BGP policies and RDMA fabric tuning.

What you will do
  • Design and deploy InfiniBand (NDR 400G / HDR) and high-speed Ethernet fabrics for GPU clusters of 1,000+ nodes
  • Configure and operate Arista, Juniper, and Mellanox/NVIDIA equipment; manage BGP, OSPF, and VXLAN overlays
  • Tune RoCE and InfiniBand transport for collective communication workloads (NCCL, UCX)
  • Maintain network automation pipelines (Ansible, Netbox, Nautobot) across all sites
  • Troubleshoot performance regressions, packet loss, and congestion during live AI training runs
What we are looking for
  • 4+ years of datacenter networking experience, including InfiniBand or 400G/800G Ethernet at scale
  • Deep familiarity with RDMA, RoCE v2, and GPU training cluster communication patterns
  • Solid command of Linux networking internals (DSCP, ECN, PFC, adaptive routing)
  • Experience with network-as-code tooling (Ansible, Terraform, Netbox) and CI/CD pipelines
  • Ability to interpret low-level diagnostics (tcpdump, perftest, ib_write_bw) and correlate with application performance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Network Engineering Lead
Cluster Network Engineering Lead

Sesterce Group • San Francisco (CA)

On-site
USD 130,000 - 160,000
Senior GPU AI Network Architect
Senior GPU AI Network Architect

Sesterce Group • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Network Engineer
Network Engineer

QumulusAI • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Equity participation
Opportunity to work on cutting-edge technology
Network Engineer - AI/HPC
Network Engineer - AI/HPC

Xai • Memphis (TN)

On-site
USD 180,000 - 240,000
Senior Technical Support Engineer, Network
Senior Technical Support Engineer, Network

NVIDIA • Germany (OH)

On-site
USD 120,000 - 180,000
Senior Technical Support Engineer, Network
Senior Technical Support Engineer, Network

NVIDIA • Town of Sweden (NY)

On-site
USD 120,000 - 160,000
Senior HPC & Infiniband Network Engineer
Senior HPC & Infiniband Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+3
Network Architect
Network Architect

TechDigital Group • California (MO)

Hybrid
USD 120,000 - 160,000
Senior Network Solution Architect – AI Fabrics
Senior Network Solution Architect – AI Fabrics

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 180,000 - 240,000