Senior Solutions Architect, Generative AI

NVIDIA Corporation

Santa Clara (CA)

Remote

USD 184,000 - 357,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation is seeking an AI Solutions Architect to design and operate large-scale AI infrastructure, optimizing GPU clusters and end-to-end workloads for leading consumer internet partners and frontier labs.

You will profile workloads, diagnose complex distributed systems, and drive Proof-of-Concepts with automation, benchmarks, and runbooks, while collaborating with engineering, product, and sales teams to win customer design wins.

Qualifications

  • BS, MS or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related field, or equivalent experience.
  • 6+ years in AI infrastructure, systems engineering, HPC, networking, or related technical roles.
  • Deep understanding of Linux, distributed computing, GPU architectures and large-scale AI clusters.
  • Hands-on experience deploying or operating high-performance GPU networks (InfiniBand, RoCE, GPUDirect RDMA).
  • Experience debugging NCCL communication and distributed performance, including topology and routing.
  • Experience with cluster schedulers (Kubernetes, Slurm), containers, and production monitoring.

Responsibilities

  • Collaborate with customers to maximize GPU utilization and workload throughput while improving reliability and reducing costs.
  • Design and optimize large-scale AI clusters across compute, networking, storage, orchestration, and observability.
  • Profile distributed workloads to identify bottlenecks across GPUs, CPUs, memory, network, and storage.
  • Diagnose complex infrastructure issues spanning InfiniBand, RoCE, cloud interconnects, NCCL, NVLink, and NVSwitch.
  • Lead proof-of-concepts and performance studies; develop benchmarking tools, automation, and runbooks.
  • Partner with engineering, product, and sales to secure design wins based on customer needs and feedback.
  • Advise on NVIDIA infrastructure technologies like DGX/HGX, NVLink, NVSwitch, NCCL, and InfiniBand.

Skills

GPU systems
AI infrastructure
Linux systems
Python
Shell scripting
distributed computing
NCCL
RDMA
Kubernetes
Slurm
Telemetry

Education

BS/MS/PhD in Computer Science/Engineering/Physics

Tools

NVIDIA DGX/HGX
InfiniBand
RoCE
GPUDirect RDMA
NCCL
Nsight Systems

Job description

NVIDIA is looking for an AI Solutions Architect with deep, hands-on experience in large-scale GPU systems. This role involves working with some of the world’s leading consumer internet companies and frontier labs building foundation models. Primary responsibilities include accelerating customer workloads, designing high-performance AI infrastructure, and leading technical engagements around NVIDIA technologies. We work with the world’s most successful technology companies, uniquely positioning you to observe and influence emerging infrastructure trends using the latest advancements. Join us in this exciting endeavour!

What You’ll Be Doing

Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs.

Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability.

Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack.

Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.

Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed.

Partnering with NVIDIA’s engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback.

What We Need To See

BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience.

6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role.

Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters.

Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using technologies such as InfiniBand, RoCE, or GPUDirect RDMA.

Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues.

Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers.

Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems.

Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.

Ways To Stand Out From The Crowd

Experience architecting and operating large-scale production GPU clusters for distributed training or inference.

Deep expertise with NVIDIA infrastructure technologies such as DGX/HGX systems, NVLink, NVSwitch, NCCL, InfiniBand, and Spectrum-X.

Hands-on experience using tools and telemetry such as NCCL tests, DCGM, Nsight Systems, fabric counters, and host- or switch-level diagnostics to isolate performance and reliability issues.

Understanding of network topology, congestion control, collective communication patterns, and their impact on distributed AI workload performance.

Experience optimizing storage and data pipelines to sustain high-throughput training and inference workloads.

We make extensive use of conferencing tools, but occasional travel (20%) is required for local on-site visits to customers and conferences. We are open to remote work.

We look forward to having you join our team!

With competitive salaries and a generous benefits package, NVIDIA is recognized as one of the technology world’s most sought-after employers.

This role offers a chance to make a broad impact at NVIDIA by advancing innovation with our consumer internet & frontier labs partners.

Are you inventive, diligent, committed, and driven? Do you enjoy tackling challenges? If so, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 3, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer.

As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry.

Learn more about NVIDIA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Solutions Architect, Generative AI
Senior Solutions Architect, Generative AI

NVIDIA • California (MO)

On-site
USD 184,000 - 356,500
Senior Solutions Architect, AI Infrastructure Enterprise ISVs
Senior Solutions Architect, AI Infrastructure Enterprise ISVs

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Solutions Architect - NVIDIA Cloud Partners
Solutions Architect - NVIDIA Cloud Partners

NVIDIA • Virginia (IL)

On-site
USD 184,000 - 357,000
Senior Software Engineer, DGX Cloud AI Infrastructure
Senior Software Engineer, DGX Cloud AI Infrastructure

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Solutions Architect - NVIDIA Cloud Partners
Solutions Architect - NVIDIA Cloud Partners

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity compensation
Benefits
Senior Solutions Architect, AI Infrastructure Enterprise ISVs
Senior Solutions Architect, AI Infrastructure Enterprise ISVs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Travel opportunities
Senior Solutions Architect, AI Compute – NPN
Senior Solutions Architect, AI Compute – NPN

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Hybrid work
Solutions Architect, Hyperscale
Solutions Architect, Hyperscale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits