Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA

Washington (Washington County)

On-site

USD 148,000 - 235,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA's First Time Deployment Team seeks an experienced systems engineer to set up, verify, and optimize AI factory environments across multi-GPU Linux clusters. You will ensure NCCL guidelines are followed and lead the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.

You'll diagnose training failures, improve observability with metrics and dashboards, develop automation in Python and Shell, study NCCL collectives, and collaborate with

Qualifications

  • Bachelor’s degree or equivalent experience in CS/Engineering/Physics or related field.
  • 6+ years of experience managing Linux-based HPC or distributed AI/ML systems.
  • Hands-on experience running AI/ML workloads on multi-GPU and multi-node clusters with NCCL.
  • Strong knowledge of AllReduce and AllToAll in ML training workflows.
  • Familiarity with PyTorch or TensorFlow for model training/inference.
  • Proficiency in Python and Shell scripting for automation.
  • Experience benchmarking and interpreting performance results.
  • Comfort working with observability data to troubleshoot complex systems.
  • Excellent written and verbal communication for cross-functional teams.

Responsibilities

  • Set up, verify, and optimize AI factory environments across multi-GPU Linux clusters.
  • Ensure configurations align with NCCL and distributed training guidelines.
  • Own execution of AI/LLM benchmarks: setup, orchestration, data collection, analysis.
  • Investigate failures, hangs, or underperformance in training jobs or benchmarks.
  • Build observability: metrics, logs, traces, dashboards to monitor workloads and health.
  • Develop automation (Python, Shell) for benchmarks, results, and regression checks.
  • Analyze communication patterns and NCCL usage for AllReduce/AllToAll workloads.
  • Recommend changes to job config, parallelism, and cluster settings for throughput/latency.
  • Collaborate with hardware, software, networking, datacenter, and product teams for readiness.
  • Contribute to documentation and readiness collateral for internal and customer teams.

Skills

Linux-based HPC
NCCL knowledge
PyTorch
TensorFlow
Python
Shell/Bash
Benchmarking
Observability data
Communication skills
Cross-functional collaboration

Education

Bachelor’s degree or equivalent

Job description

The First Time Deployment Team owns first-time execution of NVIDIA's latest products and systems; gathering install and bring-up evidence, operationalizing the validation process, documenting blockers and finding solutions to launch AI Factories at scale. Our results are spread across NVIDIA so we can succeed at scale.

What You Will Be Doing
  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
  • Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.
  • Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.
  • Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.
  • Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
  • Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks.
  • Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
  • Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
  • Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.
  • Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.
What We Need To See
  • Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.
  • More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.
  • Hands‑on experience running AI/ML workloads on multi‑GPU and/or multi-node clusters, with practical knowledge of NCCL.
  • Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training.
  • Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow.
  • Proficiency with Python and Shell/Bash for scripting, automation, and tooling.
  • Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).
  • Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.
  • Strong communication skills and the ability to work effectively with cross‑functional teams.
Ways To Stand Out From The Crowd
  • Experience with AI factory or large‑scale AI infrastructure build, deployment, or operations.
  • Background in HPC performance engineering, SRE, or systems performance analysis for GPU‑accelerated environments.
  • Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.
  • Experience building automation and CI‑style pipelines for running and validating benchmarks at scale.
  • Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data‑driven decisions.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 18, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Town of Texas (WI)

On-site
USD 148,000 - 236,000
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 236,000
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

Socket.dev • North Carolina

Hybrid
USD 152,000 - 288,000
Equity
Benefits