Senior AI Factory Architect: HPC & GPU Orchestration

NVIDIA Gruppe

California (MO)

On-site

USD 152,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking an experienced AI systems engineer to optimize and operate AI factories. You will set up multi-GPU, multi-node Linux clusters, validate NCCL configurations, run benchmarking and provide actionable improvements.

You will collaborate with hardware, software, networking, and product teams, contribute to documentation, and deliver readiness materials for internal and customer-facing teams.

Qualifications

  • Bachelor's degree or equivalent in Computer Science, Mathematics, Engineering, Physics, or related field.
  • 5+ years of experience managing Linux-based systems in HPC, distributed systems, or AI/ML environments.
  • Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, including some exposure to NCCL.
  • Practical knowledge of collective communication patterns like AllReduce and AllToAll, and their application in ML/LLM training.
  • Skilled in Python and Shell/Bash for scripting, automation, and tooling.

Responsibilities

  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
  • Validate configurations against guidelines for NCCL, collectives, and distributed training frameworks.
  • Run key AI/LLM benchmarks — setup, orchestration, result collection, and analysis.
  • Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
  • Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
  • Build automation using Python and Shell for conducting benchmarks, retrieving results, and completing regression checks.
  • Analyze communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
  • Help identify and recommend improvements to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
  • Work closely with hardware, software, networking, and product teams to prepare AI factories for customer use.
  • Contribute to documentation and readiness materials for internal and customer-facing teams.

Skills

Python
Shell/Bash
Automation
Communication
Cross-functional teamwork

Education

Bachelor's degree

Tools

NCCL
CI pipelines

Job description

NVIDIA is seeking an experienced AI systems engineer to optimize and operate AI factories. You will set up multi-GPU, multi-node Linux clusters, validate NCCL configurations, run benchmarking and provide actionable improvements.

You will collaborate with hardware, software, networking, and product teams, contribute to documentation, and deliver readiness materials for internal and customer-facing teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Factory Deployment Architect (Multi-GPU)
Senior AI Factory Deployment Architect (Multi-GPU)

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 236,000
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Architect - HPC and NCCL Benchmarks
Senior AI Factory Architect - HPC and NCCL Benchmarks

NVIDIA • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Architect - Multi-GPU Benchmarking
Senior AI Factory Architect - Multi-GPU Benchmarking

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Factory Deployment Architect
Senior AI Factory Deployment Architect

NVIDIA • Washington

On-site
USD 148,000 - 236,000
Equity
Benefits
Senior AI Factory Architect - Multi-GPU NCCL Expert
Senior AI Factory Architect - Multi-GPU NCCL Expert

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior AI/HPC Solutions Architect – Linux & Networking
Senior AI/HPC Solutions Architect – Linux & Networking

Nvidia Corporation • Santa Clara (CA)

On-site
USD 148,000 - 236,000
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits