Senior AI Network Engineer - AI Infrastructure

Hamilton Barnes Associates Limited

United States

On-site

USD 220,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual bonus
Equity opportunities
Flexible working arrangements
Career progression

Job summary

Hamilton Barnes Associates Limited is seeking a Senior AI Network Engineer to design, deploy, and optimize networking for large-scale GPU clusters driving AI workloads. You will build low-latency Ethernet/InfiniBand fabrics, configure NVIDIA Spectrum switches, and implement automation with Python, Ansible, and IaC to boost performance and reliability across HPC environments.

Collaborate with platform, storage, Kubernetes, and infra teams to maximize cluster performance, reliability, and scalable

Qualifications

  • 5+ years in Network Engineering for large-scale data centres, cloud, or AI infrastructure.
  • Deep knowledge of spine-leaf architectures and data centre networks.
  • Hands-on with NVIDIA Spectrum switches and Cumulus Linux.
  • Experience with Ethernet fabrics for AI/HPC workloads and modern routing.
  • Scripting in Python/Bash and IaC, plus automation/telemetry experience.

Responsibilities

  • Design, deploy, and operate high-performance AI networking for large GPU clusters.
  • Build low-latency Ethernet/InfiniBand fabrics for distributed AI training/inference.
  • Configure NVIDIA Spectrum, Cumulus Linux, and spine-leaf architectures.
  • Optimize performance with RDMA/RoCE/GPUDirect RDMA and NCCL.
  • Collaborate with platform/storage/Kubernetes teams to maximize cluster performance.

Skills

Spine-leaf networking
NVIDIA Spectrum switches
Cumulus Linux
RDMA / RoCE / InfiniBand
BGP / EVPN-VXLAN / MLAG
Linux systems / scripting
Infrastructure-as-Code
Network telemetry
GPU infrastructure

Tools

Python
Bash
Ansible
Kubernetes
Git

Job description

Looking for a role with plenty of growth opportunities?

Join one of North America's fastest-growing AI infrastructure providers, delivering large-scale GPU cloud platforms for AI training, fine-tuning, and inference workloads. Backed by significant investment and advanced infrastructure expertise, the organization builds high-performance AI environments powered by NVIDIA GPUs, high-speed networking, storage, and Kubernetes.

This company is seeking a Senior AI Network Engineer to design, deploy, and optimize networking infrastructure for large-scale GPU clusters. The role offers the opportunity to solve complex HPC and AI networking challenges while supporting low-latency, high-bandwidth platforms built for next-generation AI workloads.

Responsibilities:
  • Design, deploy, and operate high-performance AI networking infrastructure supporting large-scale GPU clusters.
  • Build and optimise low‑latency Ethernet and InfiniBand fabrics for distributed AI training and inference workloads.
  • Configure and support NVIDIA Spectrum switches, Cumulus Linux, and modern spine‑leaf network architectures.
  • Optimise network performance for RDMA, RoCE, GPUDirect RDMA, NCCL, and large‑scale distributed training.
  • Collaborate with platform, storage, Kubernetes, and infrastructure teams to maximise cluster performance and reliability.
  • Develop network automation using Infrastructure‑as‑Code, CI/CD, and configuration management tools.
  • Implement monitoring, telemetry, and observability across AI networking environments.
  • Troubleshoot complex networking, hardware, and distributed systems issues across production GPU infrastructure.
  • Support capacity planning, network scaling, and future infrastructure expansion.
Skills / Must Have:
  • 5+ years of experience in Network Engineering supporting large-scale data centre, cloud, HPC, or AI infrastructure environments.
  • Strong knowledge of spine‑leaf networking and large‑scale data centre architectures.
  • Hands‑on experience with NVIDIA Spectrum switches and Cumulus Linux.
  • Deep understanding of Ethernet fabrics supporting AI and HPC workloads.
  • Experience with RoCE, RDMA, InfiniBand, BGP, EVPN‑VXLAN, MLAG, and modern routing protocols.
  • Experience supporting GPU infrastructure and distributed AI training environments.
  • Strong Linux systems knowledge and scripting experience using Python, Bash, or Ansible.
  • Experience with automation, Infrastructure‑as‑Code, and network telemetry platforms.
Desirable Skills:
  • Experience supporting NVIDIA H100, H200, B200, or Blackwell GPU deployments.
  • Knowledge of NCCL, CUDA networking optimisation, GPUDirect RDMA, and distributed AI workloads.
  • Experience with Kubernetes networking and cloud‑native infrastructure.
  • Familiarity with storage networking technologies including VAST, Weka, or BeeGFS.
  • Background working within hyperscalers, GPU cloud providers, AI infrastructure companies, or HPC environments.
  • Experience deploying multi‑thousand GPU clusters.
Benefits:
  • Competitive salary with annual bonus and equity opportunities.
  • Opportunity to help build one of North America's fastest-growing AI infrastructure platforms.
  • Work with cutting‑edge NVIDIA GPU technology and hyperscale networking environments.
  • High-impact engineering role with significant technical ownership.
  • Collaborative engineering culture with minimal bureaucracy.
  • Flexible working arrangements and excellent career progression.
Salary:
  • $220,000 – $350,000 Base Salary
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior Solutions Architect, AI Infrastructure, Senior Solutions Architect, AI Infrastructure
Senior Solutions Architect, AI Infrastructure, Senior Solutions Architect, AI Infrastructure

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equities and benefits
Remote work options
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior Solution Engineer, Networking
Senior Solution Engineer, Networking

NVIDIA • Westford (MA)

On-site
USD 168,000 - 322,000
Equity
Benefits package
Senior Solution Engineer, Networking
Senior Solution Engineer, Networking

NVIDIA • Seattle (WA)

On-site
USD 168,000 - 322,000