Network Engineer

asobbi

United States

On-site

USD 160,000 - 190,000

Full time

37 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

asobbi is seeking a Network Engineer to design, build, and operate the network fabric behind a GPU cloud. This hands-on role spans multi-site data centers, RDMA networks, and edge to leaf-spine fabrics with automation-first design.

You will work across Juniper edge routing, NVIDIA/Mellanox Cumulus Linux, and Arista environments, emphasizing automations in Python, Ansible, and Terraform while ensuring robust monitoring and security posture for SOC 2/HIPAA compliance.

Qualifications

  • Experience designing, operating, and troubleshooting production network infrastructure at scale.
  • Deep understanding of TCP/IP, BGP, VLANs, VXLAN/EVPN, and modern data-center fabric design.
  • Hands-on HPC/AI networking with RDMA, InfiniBand and/or RoCE, and tuning for performance.
  • Production experience with NVIDIA/Mellanox/Cumulus Linux and/or Arista EOS in leaf-spine.
  • Strong Juniper routing and firewall experience (Junos).
  • Fluency with network automation and IaC: Python, Ansible, Terraform, Bash.
  • Packet-level (Wireshark/tcpdump) and metrics-based (Prometheus, Grafana) troubleshooting.

Responsibilities

  • Design, deployment, and tuning of high-speed network infrastructure for GPU training and inference.
  • Manage RDMA over InfiniBand and RoCEv2 with congestion control and multi-tenant isolation.
  • Edge routing, firewalling, and multi-homed connectivity on Juniper; BGP and upstream relationships.
  • Leaf-spine fabrics using NVIDIA/Mellanox Cumulus Linux and Arista in multi-site data centers.
  • Develop network automation and IaC to codify configs and fabric buildout (Python, Ansible, Terraform).
  • Support observability and troubleshooting with Prometheus, Grafana, Checkmk, Wireshark, tcpdump.
  • Collaborate on SOC 2 and HIPAA-compliant security posture; implement WireGuard VPN and related controls.

Skills

Networking at scale
BGP
VXLAN/EVPN
InfiniBand/RoCE
Juniper/Junos
Python/Ansible/Terraform
Wireshark/tcpdump

Tools

NVIDIA Mellanox hardware
Cumulus Linux
Arista EOS
Prometheus
Grafana

Job description

Location:

Remote (US)

Base:

$160,000-$190,000 + Bonus + Equity (RSUs - 4-year vest)

The Client

I'm working with a neocloud building purpose-built GPU infrastructure for AI workloads. They run large-scale clusters powering training and inference for some of the most demanding AI customers in the market.

The role

They're hiring a Network Engineer to design, build, and operate the network fabric behind their GPU cloud. This is a hands-on engineering role at the centre of a fast-moving buildout: multi-site data centres, new NVIDIA and AMD GPU clusters, and a greenfield cloud platform with automation-first network design.

You'll work across the whole stack from Juniper edge routing and firewalling, to leaf-spine fabrics on NVIDIA/Mellanox Cumulus and Arista, to the RDMA fabrics (InfiniBand and RoCE) that make large-scale training and inference possible. You'll spend as much time in automation and code as on the CLI, and you'll make trips to the data centres to help bring new capacity online.

What you'll own
  • Design, configuration, deployment, and tuning of high-speed network infrastructure for GPU training and inference across our clusters
  • The GPU fabric itself: RDMA over InfiniBand and RoCEv2 - congestion control, PFC/ECN tuning, and Pkey/partitioning for multi-tenant isolation
  • Edge routing, firewalling, and multi-homed connectivity on Juniper: BGP, public IP space, and upstream carrier relationships
  • Leaf-spine data centre fabrics on NVIDIA/Mellanox Cumulus Linux, and Arista
  • Network automation and infrastructure-as-code: codifying device configs, fabric buildout, and validation in Python, Ansible, and Terraform rather than keeping it in people's heads
  • Observability and troubleshooting across Prometheus, Grafana, Checkmk, Wireshark, and tcpdump, with a clear point of view on what good monitoring looks like
  • Network security posture alongside our security engineers: firewall policy, IDS/IPS, WireGuard VPN, and controls that hold up to SOC 2 and HIPAA
  • Data centre design: structured fibre and copper cabling, and physical-layer troubleshooting during new site and cluster buildouts
What they're looking for
Required
  • Significant experience designing, operating, and troubleshooting production network infrastructure at scale
  • Deep understanding of TCP/IP, BGP, VLANs, VXLAN/EVPN, and modern data centre fabric design
  • Hands-on high-performance networking for HPC or AI/ML: RDMA, InfiniBand and/or RoCE, and the tuning that makes them perform
  • Production experience with NVIDIA/Mellanox Cumulus Linux and/or Arista EOS in leaf-spine topologies
  • Strong Juniper routing and firewall experience (Junos)
  • Fluency with network automation and IaC: Python, Ansible, Terraform, and Bash
  • Comfortable with packet-level troubleshooting (Wireshark, tcpdump) and metrics-based troubleshooting (Prometheus, Grafana)
Non-essential but nice to have:
  • Building or operating GPU/HPC network fabrics at scale (ConnectX, Bluefield DPU, Spectrum, Quantum, UFM, or equivalent)
  • Familiarity with SDN and network-automation platforms (e.g. Netris or comparable)
  • OpenStack/OVN networking, or standing up SDN on a greenfield cloud platform
  • Working knowledge of NVIDIA GPU platforms (H100, B200) and their InfiniBand scale-out networking
  • Exposure to AMD GPU platforms (MI-series) and their Ethernet scale-out networking
  • Zero-trust network design and WireGuard VPN
  • Operating under SOC 2 and/or HIPAA
  • MAAS or bare-metal provisioning
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Network Engineer
Network Engineer

QumulusAI • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Equity participation
Opportunity to work on cutting-edge technology
Network Architect
Network Architect

TechDigital Group • Santa Clara (CA)

Hybrid
USD 120,000 - 160,000
Network Engineer
Network Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 250,000 - 320,000
GPU Network Engineer
GPU Network Engineer

CyberCoders • Santa Clara (CA)

Hybrid
USD 200,000 - 250,000
Health Benefits
401k
Relocation assistance
Senior Network Engineer
Senior Network Engineer

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 140,000 - 190,000
Senior Network Engineer - DGX Cloud
Senior Network Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 265,000
Senior Network Engineer - DGX Cloud
Senior Network Engineer - DGX Cloud

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Network Solutions Architect
Network Solutions Architect

ManpowerGroup Global, Inc. • United States

On-site
USD 120,000 - 180,000
Medical and Prescription Drug Plans
Dental Plan
Vision Plan
+8