Generative AI Solutions Architect — Remote & Equity

NVIDIA

Washington (Washington County)

On-site

USD 184,000 - 356,500

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Remote work option

Job summary

NVIDIA seeks an AI Solutions Architect in the United States to design, deploy, and optimize large-scale AI infrastructure and GPU clusters. You will lead technical engagements, accelerate customer workloads, and collaborate with engineering and sales to deliver cutting-edge NVIDIA solutions.

You will work with InfiniBand, RoCE, NCCL, and NVLink technologies, profiling workloads, improving throughput, and ensuring reliability across on-prem and cloud environments.

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience.
  • 6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role.
  • Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters.
  • Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using InfiniBand, RoCE, or GPUDirect RDMA.
  • Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues.
  • Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers.
  • Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems.
  • Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.

Responsibilities

  • Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs.
  • Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability.
  • Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack.
  • Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.
  • Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed.
  • Partnering with NVIDIA’s engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback.

Skills

Linux systems
Distributed computing
GPU architectures
Python
Shell scripting
NCCL
Kubernetes
Slurm

Education

BS/MS/PhD in CS/Engineering/Physics/Math or equivalent

Tools

NCCL
GPUDirect RDMA
InfiniBand
RoCE
Kubernetes
Slurm

Job description

NVIDIA seeks an AI Solutions Architect in the United States to design, deploy, and optimize large-scale AI infrastructure and GPU clusters. You will lead technical engagements, accelerate customer workloads, and collaborate with engineering and sales to deliver cutting-edge NVIDIA solutions.

You will work with InfiniBand, RoCE, NCCL, and NVLink technologies, profiling workloads, improving throughput, and ensuring reliability across on-prem and cloud environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GenAI Solutions Architect - Remote & GPU Infrastructure
Senior GenAI Solutions Architect - Remote & GPU Infrastructure

NVIDIA • California (MO)

Hybrid
USD 184,000 - 356,500
GenAI Solutions Architect — Enterprise-Scale AI (Equity)
GenAI Solutions Architect — Enterprise-Scale AI (Equity)

AIToolboard • Herndon (VA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior Generative AI Infrastructure Architect
Senior Generative AI Infrastructure Architect

NVIDIA Corporation • Santa Clara (CA)

Remote
USD 184,000 - 357,000
Equity
Benefits
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity and benefits
Senior AI/GPU Solutions Architect — Cloud & GenAI
Senior AI/GPU Solutions Architect — Cloud & GenAI

NVIDIA • Virginia (MN)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infra Architect for Enterprise GPU Clusters
Senior AI Infra Architect for Enterprise GPU Clusters

NVIDIA • New York (NY)

On-site
USD 184,000 - 287,500
Equity
Benefits
AI Solutions Architect - HPC/ML Leader (Equity)
AI Solutions Architect - HPC/ML Leader (Equity)

NVIDIA Corporation • Town of Texas (WI)

Hybrid
USD 184,000 - 357,000
Solutions Architecture Lead - AI, GPU Data Center (Remote)
Solutions Architecture Lead - AI, GPU Data Center (Remote)

NVIDIA • Colorado

On-site
USD 224,000 - 357,000
Senior GPU AI Solutions Architect — GenAI & HPC
Senior GPU AI Solutions Architect — GenAI & HPC

NVIDIA • New York (NY)

On-site
USD 184,000 - 288,000
Equity
Benefits
Lead AI Infra Architect for Enterprise GPU Clusters
Lead AI Infra Architect for Enterprise GPU Clusters

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 287,500