AI Network Engineer

AMSYS Innovative Solutions, LLC

Houston (TX)

On-site

USD 150,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

This hands-on, customer-facing role requires collaboration with product engineering and field teams to translate designs into field-ready solutions and deliver scalable, high-density deployments.

Qualifications

  • 10+ years designing and deploying network infrastructure for AI, HPC, or large GPU clusters.
  • Deep hands-on InfiniBand experience: rail-optimized topologies, Adaptive Routing, SHARP, congestion control.
  • RoCEv2/RDMA tuning at scale, plus strong BGP, EVPN/VXLAN, and multi-tenant design.
  • Experience with BlueField DPUs and DOCA.
  • A track record of writing architecture docs and runbooks that engineers actually use.

Responsibilities

  • Design reference architectures for AI training, inference, and HPC environments.
  • Create deployment runbooks and performance baselines for scalable GPU clusters.
  • Collaborate with field teams to translate designs into field-ready solutions.
  • Build customer-facing documentation and playbooks to support delivery and pre-sales.

Skills

AI network design
GPU cluster networking
InfiniBand
BGP
EVPN/VXLAN
ROCEv2/RDMA tuning
BlueField DPUs
Architecture docs & runbooks

Tools

Ansible
Terraform
Python

Job description

We are looking for a Sr AI Network Engineer to design and validate the fabrics behind large-scale GPU clusters for a global technology leader's AI and Hybrid Cloud engineering team. You'll build reference architectures, deployment runbooks, and performance baselines for AI training, inference, and HPC environments using NVIDIA Spectrum-X Ethernet, Quantum InfiniBand, and BlueField-3 DPUs. Your designs become the playbook that pre-sales, professional services, and delivery teams take into the field.

This is a hands-on, customer-facing role, working closely with product engineering on high-density, liquid-cooled rack deployments.

What you bring
  • 10+ years designing and deploying network infrastructure for AI, HPC, or large GPU clusters
  • Deep hands-on InfiniBand experience: rail-optimized topologies, Adaptive Routing, SHARP, and congestion control
  • RoCEv2/RDMA tuning at scale, plus strong BGP, EVPN/VXLAN, and multi-tenant design
  • Experience with BlueField DPUs and DOCA
  • A track record of writing architecture docs and runbooks that engineers actually use
Nice to have
  • NVIDIA networking or InfiniBand certification, or CCIE/CCNP Data Center
  • Automation with Ansible, Terraform, or Python
  • NeoCloud or managed service provider background
  • Exposure to FedRAMP or SOC 2 environments
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Network Engineer (AI Infrastructure) - Hosting
Network Engineer (AI Infrastructure) - Hosting

Hamilton Barnes • San Francisco (CA)

Remote
USD 191,000 - 259,000
Stock options
Remote working options and allowance
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Network Engineer
Network Engineer

asobbi • United States

On-site
USD 160,000 - 190,000
Global Data-Center Network Architect for AI/GPU HPC
Global Data-Center Network Architect for AI/GPU HPC

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 208,000 - 397,000
Equity
Benefits
Hybrid work model
Network Engineer (AI GPU Cluster Operations)
Network Engineer (AI GPU Cluster Operations)

Aquila Hash, Inc. • Buffalo (NY)

On-site
USD 110,000 - 160,000
Senior AI Networking Solutions Engineer
Senior AI Networking Solutions Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 322,000
Equity
Benefits
Network Solutions Architect
Network Solutions Architect

ManpowerGroup Global, Inc. • United States

On-site
USD 120,000 - 180,000
Medical and Prescription Drug Plans
Dental Plan
Vision Plan
+8
Network Engineer
Network Engineer

QumulusAI • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Equity participation
Opportunity to work on cutting-edge technology
Senior Network Engineer – AI GPU Infra (400G, Remote)
Senior Network Engineer – AI GPU Infra (400G, Remote)

Hamilton Barnes • San Francisco (CA)

Remote
USD 191,000 - 259,000
Stock options
Remote working options and allowance