Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA

Durham (NC)

On-site

USD 152,000 - 287,500

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Senior Solutions Architect to join the Infrastructure Specialists team in Durham, NC. You will design, implement, and validate AI factories, optimizing AI/LLM workloads on Linux GPU clusters and contributing to observability and automation efforts.

You will mentor under senior architects, collaborate across teams, and help prepare customer-ready AI factories by validating hardware and software for current AI applications.

Qualifications

  • Bachelor’s degree or equivalent in a relevant field.
  • 5+ years of experience managing Linux-based systems in HPC, distributed systems, or AI/ML environments.
  • Hands-on experience running AI/ML workloads on multi-GPU and multi-node clusters, with exposure to NCCL.
  • Practical knowledge of collective communication patterns like AllReduce and AllToAll in ML/LLM training.
  • Proficiency in Python and Shell/Bash for scripting, automation, and tooling.
  • Strong communication skills and ability to work with cross-functional teams.

Responsibilities

  • Set up, verify, and optimize AI factory environments on multi-GPU, multi-node Linux clusters.
  • Validate configurations for NCCL, collectives, and distributed training frameworks.
  • Run AI/LLM benchmarks: setup, orchestration, result collection, analysis.
  • Troubleshoot training jobs or benchmarks that fail or underperform.
  • Build observability: metrics, logs, traces, dashboards for workload understanding.
  • Automate benchmarks and regression checks using Python and Shell.
  • Analyze NCCL usage and collectives to improve throughput and latency.
  • Suggest improvements to job configuration, parallelism, and cluster settings.
  • Collaborate with hardware, software, networking, and product teams to prepare AI factories for customers.
  • Contribute to internal and customer-facing readiness materials.

Skills

Python
Shell/Bash
Communication

Education

Bachelor's degree or equivalent (CS/Engineering/Math/Physics)

Tools

NCCL

Job description

We are in search of a curious and motivated Senior Solutions Architect to join our NVIDIA Infrastructure Specialists team. In this capacity, you’ll support the creation, implementation, and verification of AI factories, focusing on running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters. You’ll engage with NCCL and collectives like AllReduce and AllToAll to boost performance and scalability, receiving mentorship from senior architects on the team.

You will apply observability and automation to improve our benchmarking and validation efforts. You will be a key contact for troubleshooting workloads and benchmarks that fail, hang, or perform poorly. Additionally, you will collaborate with various NVIDIA teams to prepare AI factories for customers, validating both hardware and software for current AI applications.

What You Will Be Doing
  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
  • Validate configurations against guidelines for NCCL, collectives, and distributed training frameworks.
  • Run key AI/LLM benchmarks - setup, orchestration, result collection, and analysis.
  • Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
  • Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
  • Build automation using Python and Shell for conducting benchmarks, retrieving results, and completing regression checks.
  • Analyze communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
  • Help identify and recommend improvements to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
  • Work closely with hardware, software, networking, and product teams to prepare AI factories for customer use.
  • Contribute to documentation and readiness materials for internal and customer-facing teams.
What We Need To See
  • Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or a related field.
  • 5+ years of experience managing Linux-based systems in HPC, distributed systems, or AI/ML environments.
  • Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, including some exposure to NCCL.
  • Practical knowledge of collective communication patterns like AllReduce and AllToAll, and their application in ML/LLM training.
  • Skilled in Python and Shell/Bash for scripting, automation, and tooling.
  • Strong communication skills and the ability to work effectively with cross-functional teams.
Ways To Stand Out From The Crowd
  • Experience benchmarking distributed systems - crafting, running, and interpreting performance benchmarks.
  • Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.
  • Familiarity with observability stacks (metrics/monitoring, logging, tracing) used for large distributed systems.
  • Experience building automation and CI-style pipelines for running and validating benchmarks at scale.
  • Demonstrated interest in using AI to solve practical problems, improve workflows, and guide data-driven decisions.

The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 4, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

, , JR2022471

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA AI • Austin (CA)

On-site
USD 124,000 - 236,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Washington

On-site
USD 148,000 - 236,000
Equity
Benefits
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Town of Texas (WI)

On-site
USD 148,000 - 236,000
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 236,000
Senior Solutions Architect, Generative AI
Senior Solutions Architect, Generative AI

NVIDIA • California (MO)

Hybrid
USD 184,000 - 357,000
Solutions Architect, AI Factory Infrastructure DevOps
Solutions Architect, AI Factory Infrastructure DevOps

NVIDIA • Austin (TX)

On-site
USD 184,000 - 288,000
Equity
Benefits