Senior AI Networking Performance Architect — Equity

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 320,000 - 489,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation, a leader in AI computing, seeks a senior software engineer to profile, analyze, and optimize AI workloads on large-scale GPU/CPU clusters for distributed Deep Learning LLM training and inference. You will focus on networking and NCCL-based communications while collaborating across hardware and software teams to drive performance improvements.

The role requires extensive experience in high-performance networking, CUDA/NVIDIA libraries, and Python/C++ tooling.

Qualifications

  • B.Sc in Computer Science or Software Engineering or equivalent experience.
  • 15+ years of experience with high-performance networking (RDMA, MPI, NCCL, SHARP).
  • Demonstrated ability in performance evaluation techniques and approaches.
  • Experience with NVIDIA GPUs and the CUDA library.
  • Knowledge of deep learning frameworks like TensorFlow or PyTorch.
  • Expertise in networking collective communication libraries such as NCCL and protocols like RoCE and RDMA.
  • Fast and self‑learning capabilities with strong analytical and problem‑solving skills.
  • Proficiency in programming languages: Python, Bash, and C++.
  • Experience with a container‑based development environment.
  • Great teammate who communicates clearly and works well with others.

Responsibilities

  • Characterizing AI workloads and deep learning models aimed at large-scale LLM training and inference on NVIDIA supercomputers.
  • The role centers on distributed systems with a focus on high-performance networking and NVIDIA communication libraries.
  • Benchmarking, profiling, and analyzing the performance to find bottlenecks and identify areas for improvement and optimizations, with a strong emphasis on networking aspects.
  • Developing PyTorch trace-based profiling, analysis, and replaying toolset to aid in benchmarking, debugging, and co‑designing network systems for LLM workloads.
  • Collaborating with multiple teams from hardware to software to provide performance analysis insights.
  • Defining performance test plans, setting performance expectations for new technologies and solutions, and working to achieve performance targets.

Job description

NVIDIA Corporation, a leader in AI computing, seeks a senior software engineer to profile, analyze, and optimize AI workloads on large-scale GPU/CPU clusters for distributed Deep Learning LLM training and inference. You will focus on networking and NCCL-based communications while collaborating across hardware and software teams to drive performance improvements.

The role requires extensive experience in high-performance networking, CUDA/NVIDIA libraries, and Python/C++ tooling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Colorado

On-site
USD 272,000 - 431,250
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 431,250
Principal Developer, AI Networking
Principal Developer, AI Networking

NVIDIA • Colorado

On-site
USD 272,000 - 431,250
Senior AI Networking Solutions Architect (Equity)
Senior AI Networking Solutions Architect (Equity)

NVIDIA • Seattle (WA)

On-site
USD 224,000 - 356,500
Equity
Benefits
Senior AI Networking Engineer - High-Performance Systems
Senior AI Networking Engineer - High-Performance Systems

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Senior Solutions Engineer - AI Networking & HPC (Equity)
Senior Solutions Engineer - AI Networking & HPC (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 322,000
Senior AI Networking Solutions Engineer
Senior AI Networking Solutions Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 322,000
Equity
Benefits
Senior AI Efficiency Engineer for GPU Clusters & Research
Senior AI Efficiency Engineer for GPU Clusters & Research

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000