Generative AI Solutions Architect — Remote & Equity

NVIDIA

Washington (Washington County)

On-site

USD 184,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Remote work option

Job summary

NVIDIA seeks an AI Solutions Architect in the United States to design, deploy, and optimize large-scale AI infrastructure and GPU clusters. You will lead technical engagements, accelerate customer workloads, and collaborate with engineering and sales to deliver cutting-edge NVIDIA solutions.

You will work with InfiniBand, RoCE, NCCL, and NVLink technologies, profiling workloads, improving throughput, and ensuring reliability across on-prem and cloud environments.

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience.
  • 6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role.
  • Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters.
  • Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using InfiniBand, RoCE, or GPUDirect RDMA.
  • Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues.
  • Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers.
  • Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems.
  • Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.

Responsibilities

  • Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs.
  • Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability.
  • Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack.
  • Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.
  • Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed.
  • Partnering with NVIDIA’s engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback.

Skills

Linux systems
Distributed computing
GPU architectures
Python
Shell scripting
NCCL
Kubernetes
Slurm

Education

BS/MS/PhD in CS/Engineering/Physics/Math or equivalent

Tools

NCCL
GPUDirect RDMA
InfiniBand
RoCE
Kubernetes
Slurm

Job description

NVIDIA seeks an AI Solutions Architect in the United States to design, deploy, and optimize large-scale AI infrastructure and GPU clusters. You will lead technical engagements, accelerate customer workloads, and collaborate with engineering and sales to deliver cutting-edge NVIDIA solutions.

You will work with InfiniBand, RoCE, NCCL, and NVLink technologies, profiling workloads, improving throughput, and ensuring reliability across on-prem and cloud environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GenAI Solutions Architect - Remote & GPU Infrastructure
Senior GenAI Solutions Architect - Remote & GPU Infrastructure

NVIDIA • California (MO)

Hybrid
USD 184,000 - 357,000
Solutions Architecture Manager - AI Leader (Remote, Equity)
Solutions Architecture Manager - AI Leader (Remote, Equity)

NVIDIA • California (MO)

On-site
USD 224,000 - 357,000
Equity
Benefits
Travel up to 20%
Senior AI Infrastructure Solutions Architect (Remote)
Senior AI Infrastructure Solutions Architect (Remote)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 287,500
Equity and benefits
GenAI Solutions Architect — Enterprise-Scale AI (Equity)
GenAI Solutions Architect — Enterprise-Scale AI (Equity)

AIToolboard • Herndon (VA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Cloud Solutions Architect — AI GPU Infra (Remote)
Cloud Solutions Architect — AI GPU Infra (Remote)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Hyperscale Data Center Solutions Architect (Remote, Equity)
Hyperscale Data Center Solutions Architect (Remote, Equity)

NVIDIA Corporation • Town of Texas (WI), Northern (KY)

Hybrid
USD 224,000 - 431,000
Senior AI Solutions Architect — GPU Cloud Infra
Senior AI Solutions Architect — GPU Cloud Infra

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Health benefits
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Solutions Architecture Manager - AI/GPU Infra (Remote)
Solutions Architecture Manager - AI/GPU Infra (Remote)

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior AI Solutions Architect – GPU Cloud & GenAI
Senior AI Solutions Architect – GPU Cloud & GenAI

NVIDIA • United States

On-site
USD 184,000 - 288,000
Equity compensation
Benefits package