Senior GenAI Solutions Architect - Remote & GPU Infrastructure

NVIDIA

California (MO)

Hybrid

USD 184,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking an AI Solutions Architect to drive high-performance AI infrastructure across large-scale GPU clusters. You will collaborate with customers to maximize GPU utilization, design scalable AI infrastructure, and lead technical engagements across NVIDIA technologies.

You will prototype and benchmark solutions, troubleshoot complex distributed systems, and partner with engineering and sales teams to deliver customer-focused design wins. Remote-ready with on-site travel as needed.

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience.
  • 6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role.
  • Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters.
  • Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using InfiniBand, RoCE, or GPUDirect RDMA.
  • Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues.
  • Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers.
  • Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems.
  • Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.

Responsibilities

  • Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs.
  • Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability.
  • Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack.
  • Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.
  • Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed.
  • Partnering with NVIDIA’s engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback.

Skills

GPU architectures
Linux systems
Distributed computing
Python scripting
NCCL debugging
Kubernetes
Slurm
RDMA
Perf profiling
Automation

Education

Bachelor/Master/PhD in Computer Science or related field

Tools

InfiniBand
NCCL
Nsight Systems

Job description

NVIDIA is seeking an AI Solutions Architect to drive high-performance AI infrastructure across large-scale GPU clusters. You will collaborate with customers to maximize GPU utilization, design scalable AI infrastructure, and lead technical engagements across NVIDIA technologies.

You will prototype and benchmark solutions, troubleshoot complex distributed systems, and partner with engineering and sales teams to deliver customer-focused design wins. Remote-ready with on-site travel as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Generative AI Solutions Architect — Remote & Equity
Generative AI Solutions Architect — Remote & Equity

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Remote work option
Solutions Architecture Manager - AI/GPU Infra (Remote)
Solutions Architecture Manager - AI/GPU Infra (Remote)

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior AI Solutions Architect — GPU Cloud Infra
Senior AI Solutions Architect — GPU Cloud Infra

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Health benefits
Cloud Solutions Architect — AI GPU Infra (Remote)
Cloud Solutions Architect — AI GPU Infra (Remote)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Solutions Architecture Lead - AI, GPU Data Center (Remote)
Solutions Architecture Lead - AI, GPU Data Center (Remote)

NVIDIA • Colorado

On-site
USD 224,000 - 357,000
Senior AI Solutions Architect – GPU Cloud & GenAI
Senior AI Solutions Architect – GPU Cloud & GenAI

NVIDIA • United States

On-site
USD 184,000 - 288,000
Equity compensation
Benefits package
Senior AI Solutions Architect — GPU Cloud
Senior AI Solutions Architect — GPU Cloud

NVIDIA • Virginia (MN)

On-site
USD 184,000 - 288,000
Senior AI Solutions Architect - GPU Cloud & Systems
Senior AI Solutions Architect - GPU Cloud & Systems

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect - Cloud GPU AI Infrastructure
Senior Solutions Architect - Cloud GPU AI Infrastructure

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GenAI Cloud Solutions Architect
Senior GenAI Cloud Solutions Architect

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits