Senior InfiniBand Network Engineer - UFM, HPC Fabric | Hybrid

Lightning AI

New York (NY)

Hybrid

USD 170,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Comprehensive Health Coverage
Meaningful Equity
401(k) matching
Unlimited PTO
Winter Break
Parental Leave
Learning allowance
Wellness stipend
Sabbatical program
Hybrid work model
In-Office Meals

Job summary

Lightning AI in New York seeks a Senior Network Engineer to design, deploy, and maintain high-performance InfiniBand fabrics for AI training and inference across GPU clusters. You will work with UFM, NCCL, and modern data center tech to ensure scalable, reliable networking.

The role emphasizes automation (Python, Ansible) and collaboration with AI platform and storage teams. Competitive compensation, equity, and benefits are offered.

Qualifications

  • 7+ years of data center networking experience.
  • 3+ years supporting NVIDIA InfiniBand environments.
  • Hands-on experience with NVIDIA UFM Enterprise.
  • Experience deploying and operating Quantum and Quantum-2 InfiniBand switches.
  • Strong Linux administration experience (Ubuntu).
  • Automation with Python and Ansible.
  • Experience with BGP, EVPN, VXLAN and spine-leaf architectures.

Responsibilities

  • Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.
  • Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
  • Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.
  • Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs.
  • Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures.
  • Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure.
  • Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication.
  • Automate network provisioning using Python, Ansible, Git, REST APIs, and IaC.
  • Monitor network health using UFM telemetry, Prometheus, Grafana, and other observability platforms.
  • Support high availability, maintenance windows, incident response, root cause analysis, and capacity planning.
  • Participate in architecture reviews and define networking standards for AI infrastructure.

Skills

InfiniBand networking
NVIDIA UFM
NVIDIA switches
Linux administration
Python automation
Ansible automation
NCCL MPI
Observability tooling

Tools

NVIDIA Quantum
NVIDIA Quantum-2
ibdiagnet
tcpdump
Wireshark
Prometheus
Grafana
UFM Enterprise
Terraform

Job description

Lightning AI in New York seeks a Senior Network Engineer to design, deploy, and maintain high-performance InfiniBand fabrics for AI training and inference across GPU clusters. You will work with UFM, NCCL, and modern data center tech to ensure scalable, reliable networking.

The role emphasizes automation (Python, Ansible) and collaboration with AI platform and storage teams. Competitive compensation, equity, and benefits are offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior InfiniBand HPC Support Engineer | Equity
Senior InfiniBand HPC Support Engineer | Equity

NVIDIA Gruppe • Westford (MA)

On-site
USD 120,000 - 207,000
Equity
Benefits
Senior Network Engineer - InfiniBand / UFM
Senior Network Engineer - InfiniBand / UFM

Lightning AI • New York (NY)

Hybrid
USD 170,000 - 210,000
Comprehensive Health Coverage
Meaningful Equity
401(k) matching
+8
Senior AI Network Engineer - Data Center & InfiniBand
Senior AI Network Engineer - Data Center & InfiniBand

Support Revolution • San Jose (CA)

On-site
USD 160,000 - 170,000
Senior InfiniBand & HPC Networking Engineer
Senior InfiniBand & HPC Networking Engineer

NVIDIA • Redmond (WA)

On-site
USD 108,000 - 173,000
Equity
Benefits
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads
HPC Cluster Engineer: Linux, InfiniBand & AI Workloads

Inflowfed • Springfield (VA)

On-site
USD 120,000 - 150,000
Senior HPC Deployment Lead: InfiniBand & AI Data Centers
Senior HPC Deployment Lead: InfiniBand & AI Data Centers

NVIDIA • Town of Texas (WI)

On-site
USD 216,000 - 397,000
Equity
Senior HPC Network Engineer | InfiniBand & Fortinet Expert
Senior HPC Network Engineer | InfiniBand & Fortinet Expert

Mirantis • United States

On-site
USD 140,000 - 190,000
Competitive compensation package
Professional development and training
Attend conferences and working groups
+1
Senior AI/HPC Solutions Architect - InfiniBand Deployments
Senior AI/HPC Solutions Architect - InfiniBand Deployments

NVIDIA • Santa Clara (CA)

On-site
USD 148,000 - 236,000
Equity
Benefits
Senior HPC Support Engineer (InfiniBand) - Equity Eligible
Senior HPC Support Engineer (InfiniBand) - Equity Eligible

NVIDIA • New York (NY)

On-site
USD 108,000 - 207,000
Equity
Comprehensive benefits