Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc

Greenhouse Software, Inc.

United States

Remote

USD 180,000 - 240,000

Full time

24 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Radian Arc is seeking a Staff Network Engineer to design, implement and operate the AI fabric, datacenter and edge networking that power our GPU cloud platform. You will shape architecture, guide cross-functional teams, and hands-on execute networking across RoCE, Ethernet, security and inter-datacenter transport.

You’ll own long-term technical direction, drive reliability and latency objectives, mentor engineers, and collaborate with platform, compute, storage, and SRE teams to embed networking

Qualifications

  • Extensive hands-on experience designing and operating large-scale datacenter networks.
  • Expert knowledge of modern networking protocols including BGP, ECMP and VRF.
  • Experience with GPU networking and AI fabric environments is preferred.

Responsibilities

  • Design and operate high-performance GPU networking fabrics for distributed AI workloads.
  • Architect RoCE fabrics optimized for training and inference.
  • Collaborate with compute, storage and SRE teams to integrate networking across platforms.

Skills

AI Fabric Networking
Datacenter Networking
RoCE
BGP/ECMP
Network Automation
Security & Edge Connectivity

Tools

VyOS
OVS/OVN
BlueField DPU
NVIDIA Mellanox platforms

Job description

Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc

Remote

Type of Contract: Full time or Contract

About Radian Arc

We’re specialists in outcome-optimized AI infrastructure - deploying, orchestrating and monetizing GPU compute where data, users and demand actually meet: inside telco networks, at the edge, and in core data centers.Not a generic AI platform. Not a consultancy. We’re the bridge between raw silicon and real-world results.

What impact you will have

Design, implement, and operate the network infrastructure powering the GPU cloud platform,including high-performance AI fabrics as well as classical datacenter networking components suchas routing, security, and external connectivity.This role spans both high-performance east-west networking for distributed AI workloads andnorth-south connectivity, security, and inter-datacenter transport.

As the first dedicated networking role in the organization, the Staff Network Engineer combinesStaff-level architectural ownership, technical direction, and cross-functional influence with hands-onexecution across design, deployment, troubleshooting, automation, and operational improvement.

The Staff Network Engineer owns the long-term technical direction and operational strategy forRadian Arc’s AI interconnect networks, designing scalable GPU fabrics and ensuring predictablelow-latency performance across distributed training and inference workloads. The role includesdesigning large-scale RoCE and Ethernet fabrics, guiding architecture decisions, and ensuringoperational excellence across global deployments, from hyperscale datacenters to smaller edgelocations.

You will collaborate closely with platform, compute, storage, observability, and operations teams toensure networking is deeply integrated into the overall infrastructure architecture. This role also actsas the senior escalation point for complex networking incidents, driving deep technical investigationsand systemic improvements that increase reliability, latency consistency, and operational maturityacross the platform.

Because this is currently the primary networking role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of technical direction, standards,cross-team influence, and long-term design, while also directly executing critical networking work that, in a larger organization, would be distributed across multiple engineers.

We expect to hire more than one person for this role so providing you have experience in designing AI Infrastructure networking solutions at scale please do not hesitate to apply if your experience does not match the full scope of the position.

What you’ll do
AI Fabric & HPC Networking
  • Design and operate high-performance GPU networking fabrics supporting distributed AI workloads.
  • Architect large-scale RoCE fabrics optimized for distributed training and inference.
  • Optimize network performance for GPU communication patterns and east-west traffic.
  • Design fabric topologies such as:
  • Implement high-performance networking technologies including:
  • High-bandwidth east-west fabrics

Collaborate with compute teams to support distributed training frameworks and GPU communication libraries.

Define reference architectures and design principles for AI fabrics so future deployments follow reusable standards rather than one-off implementations.

Evaluate architectural trade-offs across performance, resilience, cost, operability, and deployment speed, and make clear recommendations to stakeholders.

Datacenter Networking
  • Design and operate Layer-2 and Layer-3 datacenter networks.
  • Implement scalable routing architectures based on BGP and ECMP.
  • Implement and maintain:
  • Maintain north-south ingress/egress routing and traffic management.
  • Define standards and reusable patterns for segmentation, routing, and overlay integration across platform deployments.
Technologies include:
  • VyOS routers
  • OVS / OVN
  • BGP / ECMP
Security & Edge Connectivity
  • Deploy and maintain north-south security infrastructure
  • Implement WAF and application-layer protections
  • Integrate security controls with platform services
Technologies include:
  • TLS termination
  • API and proxy gateway protection
Inter-Datacenter Networking
  • Design and operate private interconnects between datacenters
  • Implement and maintain dark fiber ring architectures
  • Operate high-capacity WAN connectivity between regions
  • Define scalable design principles for backbone evolution, inter-site routing, redundancy, and failure-domain isolation
Technologies include:
  • BGP inter-site routing
  • Redundant fiber ring architectures
  • 100–400G optical transport
  • Spectrum-XGS
  • Lead end-to-end engineering delivery of networking infrastructure, from design and labvalidation to production deployment
  • Validate network BOMs together with procurement and deployment teams
  • Provide detailed input into datacenter layouts and rack elevations
  • Drive capacity planning, performance modeling, and scaling strategies
  • Ensure network changes are executed safely with minimal customer impact
  • Act as both the architectural owner and the practical execution lead for critical network initiatives during the build-out phase of the networking function
  • Establish deployment standards, validation criteria, rollback approaches, and acceptance patterns that future engineers and teams can reuse
Operational Excellence & Reliability
  • Own operational performance and reliability of networking infrastructure
  • Lifecycle management
  • Improve day-2 operations through automation and operational tooling
  • Lead incident response and root-cause analysis for major network events
  • Define and track SLAs, SLOs, and reliability metrics
  • Translate major incidents and operational pain points into durable standards, design changes, and long-term architectural improvements
  • Establish measurable benchmarks for reliability, latency consistency, operability, and recovery behavior across network deployments.
Cross-Functional Collaboration
  • Work closely with infrastructure, platform, SRE, compute, storage, observability, and datacenter operations teams.
  • Provide technical leadership across infrastructure initiatives.
  • Communicate architectural decisions, trade-offs, and risks clearly to stakeholders.
  • Influence the long-term platform networking roadmap and architecture.
  • Act as the primary networking design authority across the organization, guiding adjacent teams on how networking constraints and capabilities should shape platform decisions.
  • Raise the technical bar by mentoring engineers in adjacent domains and helping build the future networking function.
Technical Stack
Datacenter Networking
  • BGP
  • EVPN / VXLAN
  • ECMP
  • VLAN / VRF
  • OVS / OVN
  • BlueField DPU
Routing & Control Plane
  • VyOS
  • BGP-based routing architectures
  • ECMP fabrics
Security
  • DDoS protection
Transport & Backbone
  • Dark fiber
  • DWDM transport
  • 100–1600G optical networking
AI Networking
  • RDMA
  • RoCE
  • GPU fabrics
  • Large-scale east-west compute networking
  • Congestion control
What you'll need
Core Experience
  • Strong hands-on experience designing and operating large-scale datacenter networks
  • Expert knowledge of modern networking protocols including:
  • Proven experience operating high-speed Ethernet networks in production environments
  • Experience operating NVIDIA / Mellanox networking platforms
  • Experience owning both architecture and direct implementation in lean or fast-scaling environments is strongly preferred
Advanced AI Fabric Networking Expertise

The candidate should have deep expertise in designing and operating networking fabrics optimizedfor large-scale GPU clusters and distributed AI workloads.

This includes a strong understanding of GPU communication patterns and the networkingrequirements of distributed training and inference systems.

  • Deep understanding of NCCL communication patterns and their impact on network topology and performance.
  • Experience tuning RoCE fabrics for large-scale GPU clusters.
  • Strong knowledge of RDMA transport behavior and failure modes.
  • Practical experience implementing and tuning PFC and ECN for congestion management.
  • Understanding of GPU collective communication patterns such as all-reduce, all-gather, broadcast, reduce-scatter, and their impact on east-west network traffic.
  • Experience designing rail-optimized GPU networking fabrics for distributed training and inference clusters.
  • Familiarity with diagnosing performance issues related to:
  • Understanding of how networking performance affects distributed AI frameworks such as PyTorch and TensorFlow.

The candidate should also be able to collaborate closely with compute platform teams to ensure thatnetworking infrastructure is optimized for distributed training, distributed inference, andhigh-throughput AI workloads.

Systems & Troubleshooting
  • Ability to debug complex cross-layer issues spanning:
Distributed application communication layers
  • Strong knowledge of networking hardware, optics, and high-speed interconnects.
  • Experience designing network observability systems.
  • Strong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues.
Automation
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Network Engineer
Network Engineer

asobbi • United States

On-site
USD 160,000 - 190,000
Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 120,000 - 190,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Network Engineer
Network Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 180,000
Network Engineer (AI GPU Cluster Operations)
Network Engineer (AI GPU Cluster Operations)

Aquila Hash, Inc. • Buffalo (NY)

On-site
USD 110,000 - 160,000
Network Architect
Network Architect

GMI Cloud, Inc • United States

On-site
USD 180,000 - 240,000
Senior Network Engineer
Senior Network Engineer

Info Way Solutions • Austin (TX)

On-site
USD 120,000 - 180,000
Staff HPC Network Architect
Staff HPC Network Architect

Lambda • United States

Hybrid
USD 180,000 - 280,000
Senior/Staff Software Engineer, Network Infrastructure
Senior/Staff Software Engineer, Network Infrastructure

Kindredventures • United States

On-site
USD 180,000 - 240,000