Senior AI Factory Deployment Architect (Multi-GPU)

NVIDIA

Virginia (MN)

On-site

USD 148,000 - 235,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a senior engineer to design and operate AI factory environments across multi-GPU, multi-node Linux clusters. You will ensure NCCL and collectives configurations align with best practices, run key AI/LLM benchmarks, and analyze results for optimization.

You will develop automation in Python and Shell, improve observability through metrics, logs, and dashboards, and work with cross-functional teams to scale deployments. A strong background in HPC and ML workloads is required.

Qualifications

  • Bachelor's degree or equivalent in CS, Engineering, Mathematics, Physics, or related field.
  • 6+ years of experience managing Linux-based HPC or AI/ML workloads.
  • Hands-on experience with AI/ML workloads on multi-GPU/multi-node clusters; practical NCCL knowledge.
  • Solid understanding of AllReduce and AllToAll in ML training.
  • Familiarity with PyTorch or TensorFlow for LL(M) workloads.
  • Proficiency in Python and Shell scripting for automation.
  • Experience benchmarking and interpreting performance metrics.
  • Comfort with observability data to troubleshoot distributed systems.
  • Strong communication and collaboration across functions.

Responsibilities

  • Set up, adjust, and verify AI factory environments on Linux clusters.
  • Configure NCCL, collectives, and distributed training workflows.
  • Execute, collect, and analyze AI/LLM benchmarks.

Skills

Linux systems
HPC
NCCL
Python
Shell/Bash
Benchmarking
Observability data
PyTorch
TensorFlow
Multi-GPU / Multi-node
AllReduce / AllToAll
Cross-functional collaboration

Education

Bachelor's degree or equivalent in CS/Engineering/Math/Physics

Tools

Python
Shell
NCCL Toolkit

Job description

NVIDIA is seeking a senior engineer to design and operate AI factory environments across multi-GPU, multi-node Linux clusters. You will ensure NCCL and collectives configurations align with best practices, run key AI/LLM benchmarks, and analyze results for optimization.

You will develop automation in Python and Shell, improve observability through metrics, logs, and dashboards, and work with cross-functional teams to scale deployments. A strong background in HPC and ML workloads is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Deployment Architect
Senior AI Factory Deployment Architect

NVIDIA • Washington

On-site
USD 148,000 - 236,000
Equity
Benefits
Senior AI Factory Architect - Multi-GPU NCCL Expert
Senior AI Factory Architect - Multi-GPU NCCL Expert

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Architect: HPC & GPU Orchestration
Senior AI Factory Architect: HPC & GPU Orchestration

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior AI Factory Architect - Multi-GPU Benchmarking
Senior AI Factory Architect - Multi-GPU Benchmarking

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Architect - HPC and NCCL Benchmarks
Senior AI Factory Architect - HPC and NCCL Benchmarks

NVIDIA • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Solutions Architect, AI Factory Deployment - NVIS
Senior Solutions Architect, AI Factory Deployment - NVIS

NVIDIA • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Compute Deployment Engineer
Senior AI Compute Deployment Engineer

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity and Benefits