Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA Corporation

Santa Clara (CA)

Hybrid

USD 148,000 - 236,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA Corporation in Santa Clara, CA seeks a Senior Solutions Architect to drive validation of AI factories from first rack power-on through customer handoff. You will run and debug AI/LLM workloads on Linux-based GPU clusters, collect evidence, and work with internal teams to ensure readiness for launch.

You’ll optimize NCCL usage, build observability, and automate benchmarks with Python and Shell. Strong collaboration and problem-solving skills are essential for scalable, high-performance AI

Qualifications

  • Bachelor’s degree or equivalent in CS, Math, Engineering or related field.
  • 5+ years of Linux-based HPC or distributed systems experience.
  • Hands-on AI/ML workloads on multi-GPU clusters with NCCL.
  • Strong knowledge of AllReduce/AllToAll patterns in ML training.

Responsibilities

  • Set up and verify AI factory environments across multi-GPU/multi-node Linux clusters.
  • Ensure configurations align with NCCL, collectives and distributed training frameworks.
  • Own execution of AI/LLM benchmarks: setup, orchestration, result collection and analysis.
  • Investigate and resolve training/benchmark issues
  • Build observability for AI factories: metrics, logs, traces, dashboards
  • Develop automation (Python, Shell) for benchmarks and regression checks
  • Examine NCCL usage and communication patterns for ML workloads
  • Recommend changes to configurations to improve throughput and scaling

Skills

Linux HPC
AI/ML workloads
NCCL
PyTorch
TensorFlow
Python
Shell/Bash
Benchmarking
Communication
Cross-functional teamwork

Education

Bachelor’s degree or equivalent in CS/Math/Engineering/Physics

Tools

NCCL
PyTorch
TensorFlow
Bash

Job description

The First Time Deployment Team owns first-time execution of NVIDIA's latest products and systems; gathering install and bring-up evidence, operationalizing the validation process, documenting blockers and finding solutions to launch AI Factories at scale. Our results are spread across NVIDIA so we can succeed at scale.We're looking for an ambitious Senior Solutions Architect to drive validation of NVIDIA AI factories from first rack power-on through customer handoff. You will be embedded in launches from the start, running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters using NCCL and collectives (AllReduce, AllToAll) to validate performance and scalability. When workloads or benchmarks fall short, you're the expert who digs in, partners with engineering, and drives resolution. You will operationalize observability and automation to accelerate validation, capture structured evidence across every bring-up milestone, and work directly with internal deployment teams and external customers to ensure AI factories are ready at launch. Your work directly enables the success of NVIDIA's first external product launches!**What You Will be Doing:*** Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.* Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.* Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.* Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.* Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.* Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks* Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.* Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.* Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.* Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.**What We Need to See:*** Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.* More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.* Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL.* Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training.* Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow.* Proficiency with Python and Shell/Bash for scripting, automation, and tooling.* Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).* Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.* Strong communication skills and the ability to work effectively with cross-functional teams.**Ways to Stand Out From the Crowd:*** Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.* Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.* Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.* Experience building automation and CI-style pipelines for running and validating benchmarks at scale.* Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.NVIDIA is widely considered one of the technology world’s most desirable employers. Some of the world's most forward-thinking and hardworking people are working for us. If you're creative and autonomous, we want to hear from you.#BuildTheAIFactoryYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD.You will also be eligible for equity and benefits.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Washington

On-site
USD 148,000 - 236,000
Equity
Benefits
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Town of Texas (WI)

On-site
USD 148,000 - 236,000
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 236,000
Manager, First Time Deployment - NVIS
Manager, First Time Deployment - NVIS

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 345,000
Equity and benefits
Senior Solutions Architect, AI Factory Observability and Visualization - NVIS
Senior Solutions Architect, AI Factory Observability and Visualization - NVIS

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Solutions Architect, AI Factory Observability and Visualization - NVIS
Senior Solutions Architect, AI Factory Observability and Visualization - NVIS

NVIDIA Corporation • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity options
Comprehensive benefits package
Senior Infrastructure Engineer - Infrastructure Security and Core Services
Senior Infrastructure Engineer - Infrastructure Security and Core Services

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 208,000 - 334,000
Senior AI Compute Engineer - NVIS
Senior AI Compute Engineer - NVIS

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 148,000 - 235,750
Senior Solutions Architect - AI Infrastructure
Senior Solutions Architect - AI Infrastructure

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Senior Solutions Architect, Ethernet Networking - NVIS
Senior Solutions Architect, Ethernet Networking - NVIS

NVIDIA Corporation • Northern (KY)

Hybrid
USD 148,000 - 288,000
Equity
Benefits