QumulusAI is hiring a Network Engineer to design, build, and operate high-performance network fabrics powering AI/ML GPU infrastructure at scale. You will work on cutting-edge leaf-spine architectures, high-speed lossless fabrics spanning Ethernet and InfiniBand, and security infrastructure that supports thousands of GPUs delivering petaflops of compute. This role is critical to ensuring our network platform delivers the performance, reliability, and security our customers demand across current and next-generation interconnect technologies.
What You Will Do
- Design, implement, and maintain leaf-spine network fabrics using Arista, Cisco, SONiC, and/or Juniper platforms in multi-datacenter environments
- Configure and optimize VXLAN EVPN overlays, BGP underlay routing (eBGP/iBGP), and MLAG/LACP link aggregation across production network infrastructure
- Deploy, manage, and maintain firewall and security infrastructure using FortiGate, Cisco, and/or Palo Alto platforms including policy management, VPN (IPSec/SSL), and threat prevention
- Operate and optimize high-speed lossless fabrics (400G, 800G, and beyond) for GPU-to-GPU communication including PFC, ECN, and DCQCN configuration across Ethernet and InfiniBand interconnects
- Develop and maintain network automation using Python, Ansible, and/or platform-native APIs (eAPI, NETCONF, REST) for configuration management and operational workflows
- Perform capacity planning, traffic analysis, and performance optimization across IP transit, peering, and customer-facing network segments
- Participate in on-call rotations, incident response, root cause analysis, and any and all engineering tasks required to maintain and advance the platform
- Provide outstanding customer service when engaging directly with customers, including network troubleshooting, technical guidance, and escalation support
- Manage IP address space (IPAM), DNS infrastructure, and out-of-band management networks
What We Are Looking For
- 5+ years of network engineering experience in datacenter, cloud, or service provider environments
- Strong hands-on experience with Arista EOS including VXLAN, EVPN, BGP, and MLAG configuration
- Proficiency with at least one additional vendor platform: SONiC, Cisco IOS-XE/NX-OS, Juniper Junos, or equivalent
- Solid experience with enterprise firewall platforms: FortiGate (FortiOS), Cisco, and/or Palo Alto (PAN-OS) including policy design, NAT, VPN, and HA configurations
- Deep understanding of L2/L3 networking: spanning tree, VLANs, OSPF, BGP (eBGP/iBGP), route policies, and prefix filtering
- Experience with network automation and programmability: Python, Ansible, Jinja2 templates, and API-driven configuration
- Familiarity with RDMA networking concepts: lossless Ethernet (RoCE), InfiniBand, PFC, ECN, and quality of service configuration
- Professional network certification preferred: Arista ACE, CCNP/CCIE, JNCIP/JNCIE, or equivalent demonstrated expertise
Nice to Have
- Experience with TCAM optimization, hardware forwarding table management, and platform-specific resource constraints
- Knowledge of IP transit purchasing, peering agreements, and BGP community-based traffic engineering
- Familiarity with network observability platforms: SNMP, streaming telemetry, NetFlow/sFlow, and tools like LibreNMS or Grafana
- Experience with out-of-band management networks, console servers, and IPMI/BMC network segmentation
- Background in GPU datacenter networking including spine-leaf design for AI/ML workloads and GPUDirect RDMA
- Experience with network security compliance frameworks and audit requirements
QumulusAI is building the next generation of AI infrastructure. We operate high-performance bare-metal GPU datacenters with cutting-edge network fabrics spanning the latest Ethernet and InfiniBand technologies, serving customers who are pushing the boundaries of artificial intelligence. Our team combines deep technical expertise with startup agility, and every engineer has a direct impact on the platform and the business.
We offer competitive compensation, equity participation, and the opportunity to work on infrastructure challenges at a scale that few companies in the world encounter. If you thrive on solving hard problems with world-class technology, we want to hear from you.