Distinguished Software Architect - Deep Learning and HPC Communications

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 320,000 - 488,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity options
Comprehensive benefits package

Job summary

NVIDIA Gruppe in Santa Clara is seeking a Distinguished Software Architect to lead the design of next-generation data center platforms. This role demands deep expertise in HPC and networking, aiming to improve GPU communication technologies.

You will research and implement innovative solutions while collaborating with internal and external teams. The position offers opportunities to impact cutting-edge technology and a competitive salary range of 320,000 to 488,750 USD based on experience.

Qualifications

  • 15+ years of relevant experience in academia or industry.
  • Expertise in computer and system architecture.
  • Flexibility to work and communicate across different teams.

Responsibilities

  • Research new communication technologies and design new features.
  • Propose innovative hardware and software solutions.
  • Drive the adoption of new communication technologies.

Skills

HPC expertise
Parallel programming models (MPI, SHMEM)
C/C++ programming fluency
Deep understanding of networking technologies (Infiniband, Ethernet)
ML/DL fundamentals knowledge

Education

PhD in Computer Science or related field

Job description

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars.

GPU Communications Libraries and Networking team
Distinguished Software Architect

We are looking for a Distinguished Software Architect to help co‑design our next generation data center platforms. DL and HPC applications have a huge compute demand already and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high‑speed interconnects (e.g., NVLink, PCIe) within a node and with high‑speed networking (e.g., Infiniband, Ethernet) across nodes. Communication performance between the GPUs directly impacts end‑to‑end application performance; the stakes are even higher at huge scales. This is an outstanding opportunity to push the limits of state‑of‑the‑art technologies and deliver platforms the world has never seen before.

What you will be doing
  • Research new communication technologies (e.g., expand the GPUDirect technology portfolio) and design new features for our communication libraries.
  • Propose innovative solutions in hardware and software for our next‑gen platforms. You will co‑design these solutions with the GPU, Networking, and SW architects and ensure seamless integration with the software stacks.
  • Inspire changes based on quantitative data from proof‑of‑concepts or detailed technical analysis/modeling.
  • Drive the adoption of new communication technologies across application verticals.
  • Keep up with the latest DL research and collaborate with diverse teams (internal and external), including DL researchers and customers.
What we need to see
  • PHD in Computer Science, Computer Engineering or related field or strong equivalent experience; 15+ years of relevant experience in academia or industry.
  • Expertise in HPC, parallel programming models (MPI, SHMEM), at least one communication runtime (MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, UCC), computer and system architecture, GPU architecture and CUDA.
  • Deep understanding of high‑performance networking aspects: network technologies (Infiniband, Ethernet), network design, topologies, debugging and performance analysis.
  • Strong knowledge in at least a few of these areas: ML/DL fundamentals and their relation to communications, parallel algorithms, fault tolerance and resiliency, competitive assessments, performance analysis and optimizations for large clusters, developing applications using DL frameworks (PyTorch, TensorFlow).
  • Programming fluency with C or C++ for systems software development.
  • Flexibility to work and communicate effectively across different HW/SW teams and time zones.
Ways to stand out from the crowd
  • Industry‑recognized leader in HPC/DL communications with a history of patents, publications, conference talks and keynotes relevant to the role.
  • Influential role in industry standards (e.g., MPI, OpenSHMEM) and open‑source software (e.g., PyTorch, UCX, Open MPI).
Compensation and Benefits

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD – 488,750 USD. You will also be eligible for equity and benefits.

Application Deadline

Applications for this job will be accepted until May 26, 2026.

Equal Opportunity Statement

NVIDIA is committed to fostering a diverse work environment and is a proud equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Software Architect - Deep Learning and HPC Communications
Distinguished Software Architect - Deep Learning and HPC Communications

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 488,750
Senior Software Architect - Deep Learning and HPC Communications
Senior Software Architect - Deep Learning and HPC Communications

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Software Architect - Deep Learning and HPC Communications
Senior Software Architect - Deep Learning and HPC Communications

NVIDIA • Westford (MA)

On-site
USD 224,000 - 357,000
Equity
Comprehensive benefits
Senior Software Architect - Deep Learning and HPC Communications
Senior Software Architect - Deep Learning and HPC Communications

NVIDIA • Austin (TX)

On-site
USD 184,000 - 288,000
Senior Software Architect - Deep Learning and HPC Communications
Senior Software Architect - Deep Learning and HPC Communications

NVIDIA AI • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits package
Senior Software Architect - Deep Learning and HPC Communications
Senior Software Architect - Deep Learning and HPC Communications

NVIDIA • Durham (NC)

On-site
USD 120,000 - 160,000
Principal Deep Learning Communication Architect
Principal Deep Learning Communication Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Competitive salary
Equity options
Generous benefits package
Principal Architect, AI Networking
Principal Architect, AI Networking

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Senior Software Architect - Deep Learning and HPC Communications
Senior Software Architect - Deep Learning and HPC Communications

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 356,500
Equity and benefits
Senior Software Engineer, NCCL
Senior Software Engineer, NCCL

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Comprehensive benefits package
Equity opportunities