Looking for a role with plenty of growth opportunities?
Join one of North America's fastest-growing AI infrastructure providers, delivering large-scale GPU cloud platforms for AI training, fine-tuning, and inference workloads. Backed by significant investment and advanced infrastructure expertise, the organization builds high-performance AI environments powered by NVIDIA GPUs, high-speed networking, storage, and Kubernetes.
This company is seeking a Senior AI Network Engineer to design, deploy, and optimize networking infrastructure for large-scale GPU clusters. The role offers the opportunity to solve complex HPC and AI networking challenges while supporting low-latency, high-bandwidth platforms built for next-generation AI workloads.
Responsibilities:
- Design, deploy, and operate high-performance AI networking infrastructure supporting large-scale GPU clusters.
- Build and optimise low‑latency Ethernet and InfiniBand fabrics for distributed AI training and inference workloads.
- Configure and support NVIDIA Spectrum switches, Cumulus Linux, and modern spine‑leaf network architectures.
- Optimise network performance for RDMA, RoCE, GPUDirect RDMA, NCCL, and large‑scale distributed training.
- Collaborate with platform, storage, Kubernetes, and infrastructure teams to maximise cluster performance and reliability.
- Develop network automation using Infrastructure‑as‑Code, CI/CD, and configuration management tools.
- Implement monitoring, telemetry, and observability across AI networking environments.
- Troubleshoot complex networking, hardware, and distributed systems issues across production GPU infrastructure.
- Support capacity planning, network scaling, and future infrastructure expansion.
Skills / Must Have:
- 5+ years of experience in Network Engineering supporting large-scale data centre, cloud, HPC, or AI infrastructure environments.
- Strong knowledge of spine‑leaf networking and large‑scale data centre architectures.
- Hands‑on experience with NVIDIA Spectrum switches and Cumulus Linux.
- Deep understanding of Ethernet fabrics supporting AI and HPC workloads.
- Experience with RoCE, RDMA, InfiniBand, BGP, EVPN‑VXLAN, MLAG, and modern routing protocols.
- Experience supporting GPU infrastructure and distributed AI training environments.
- Strong Linux systems knowledge and scripting experience using Python, Bash, or Ansible.
- Experience with automation, Infrastructure‑as‑Code, and network telemetry platforms.
Desirable Skills:
- Experience supporting NVIDIA H100, H200, B200, or Blackwell GPU deployments.
- Knowledge of NCCL, CUDA networking optimisation, GPUDirect RDMA, and distributed AI workloads.
- Experience with Kubernetes networking and cloud‑native infrastructure.
- Familiarity with storage networking technologies including VAST, Weka, or BeeGFS.
- Background working within hyperscalers, GPU cloud providers, AI infrastructure companies, or HPC environments.
- Experience deploying multi‑thousand GPU clusters.
Benefits:
- Competitive salary with annual bonus and equity opportunities.
- Opportunity to help build one of North America's fastest-growing AI infrastructure platforms.
- Work with cutting‑edge NVIDIA GPU technology and hyperscale networking environments.
- High-impact engineering role with significant technical ownership.
- Collaborative engineering culture with minimal bureaucracy.
- Flexible working arrangements and excellent career progression.
Salary:
- $220,000 – $350,000 Base Salary