Senior AI Factory Deployment Architect

NVIDIA

Washington (Washington County)

On-site

USD 148,000 - 235,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA's First Time Deployment Team seeks an experienced systems engineer to set up, verify, and optimize AI factory environments across multi-GPU Linux clusters. You will ensure NCCL guidelines are followed and lead the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.

You'll diagnose training failures, improve observability with metrics and dashboards, develop automation in Python and Shell, study NCCL collectives, and collaborate with

Qualifications

  • Bachelor’s degree or equivalent experience in CS/Engineering/Physics or related field.
  • 6+ years of experience managing Linux-based HPC or distributed AI/ML systems.
  • Hands-on experience running AI/ML workloads on multi-GPU and multi-node clusters with NCCL.
  • Strong knowledge of AllReduce and AllToAll in ML training workflows.
  • Familiarity with PyTorch or TensorFlow for model training/inference.
  • Proficiency in Python and Shell scripting for automation.
  • Experience benchmarking and interpreting performance results.
  • Comfort working with observability data to troubleshoot complex systems.
  • Excellent written and verbal communication for cross-functional teams.

Responsibilities

  • Set up, verify, and optimize AI factory environments across multi-GPU Linux clusters.
  • Ensure configurations align with NCCL and distributed training guidelines.
  • Own execution of AI/LLM benchmarks: setup, orchestration, data collection, analysis.
  • Investigate failures, hangs, or underperformance in training jobs or benchmarks.
  • Build observability: metrics, logs, traces, dashboards to monitor workloads and health.
  • Develop automation (Python, Shell) for benchmarks, results, and regression checks.
  • Analyze communication patterns and NCCL usage for AllReduce/AllToAll workloads.
  • Recommend changes to job config, parallelism, and cluster settings for throughput/latency.
  • Collaborate with hardware, software, networking, datacenter, and product teams for readiness.
  • Contribute to documentation and readiness collateral for internal and customer teams.

Skills

Linux-based HPC
NCCL knowledge
PyTorch
TensorFlow
Python
Shell/Bash
Benchmarking
Observability data
Communication skills
Cross-functional collaboration

Education

Bachelor’s degree or equivalent

Job description

NVIDIA's First Time Deployment Team seeks an experienced systems engineer to set up, verify, and optimize AI factory environments across multi-GPU Linux clusters. You will ensure NCCL guidelines are followed and lead the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.

You'll diagnose training failures, improve observability with metrics and dashboards, develop automation in Python and Shell, study NCCL collectives, and collaborate with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Factory Deployment Architect (Multi-GPU)
Senior AI Factory Deployment Architect (Multi-GPU)

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 236,000
Senior AI Factory Architect: GPU Clusters & Benchmarks
Senior AI Factory Architect: GPU Clusters & Benchmarks

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Factory Architect: GPUs, NCCL, Automation, Equity
Senior AI Factory Architect: GPUs, NCCL, Automation, Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Factory Architect: HPC & GPU Orchestration
Senior AI Factory Architect: HPC & GPU Orchestration

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Architect - Multi-GPU NCCL Expert
Senior AI Factory Architect - Multi-GPU NCCL Expert

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Architect - HPC and NCCL Benchmarks
Senior AI Factory Architect - HPC and NCCL Benchmarks

NVIDIA • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Architect - Multi-GPU Benchmarking
Senior AI Factory Architect - Multi-GPU Benchmarking

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Washington

On-site
USD 148,000 - 236,000
Equity
Benefits
Senior Solutions Architect, First Time Deployment Validation - NVIS
Senior Solutions Architect, First Time Deployment Validation - NVIS

NVIDIA • Town of Texas (WI)

On-site
USD 148,000 - 236,000