Senior AI Systems Tools Engineer - GPU Clusters

NVIDIA Gruppe

Santa Clara (CA)

Hybrid

USD 184,000 - 357,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking engineers to develop tools for AI researchers and software/hardware teams running AI workloads on GPU clusters. You will build profiling, analysis, and debugging tools at scale.

The role emphasizes strong C++/Python skills, experience with PyTorch/TensorFlow, and collaboration with hardware architects to deliver impactful features and improvements. Hybrid work is supported in our global setup.

Qualifications

  • BS+ in Computer Science or related field (or equivalent experience) and 6+ years of software development.
  • Strong skills in design, coding (C++ and Python), analytical thinking and debugging.
  • Knowledge of PyTorch and TensorFlow; distributed training and inference.

Responsibilities

  • Build internal profiling and analysis tools for AI workloads at large scale.
  • Build debugging tools for memory, networking, and other common problems.
  • Create benchmarking and simulation technologies for AI systems or GPU clusters.
  • Partner with HW architects to propose new features or improve existing ones with real-world use cases.

Skills

C++
Python
Debugging
Distributed training
Problem solving

Education

BS in Computer Science or related

Tools

PyTorch
TensorFlow
Slurm
Kubernetes
CUDA
NCCL
Linux

Job description

NVIDIA is seeking engineers to develop tools for AI researchers and software/hardware teams running AI workloads on GPU clusters. You will build profiling, analysis, and debugging tools at scale.

The role emphasizes strong C++/Python skills, experience with PyTorch/TensorFlow, and collaboration with hardware architects to deliver impactful features and improvements. Hybrid work is supported in our global setup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Performance Tools Engineer (GPU Profiling)
Senior AI Performance Tools Engineer (GPU Profiling)

NVIDIA • United States

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer — GPU Clusters
Senior AI Infrastructure Engineer — GPU Clusters

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Performance & Efficiency Engineer - GPU Clusters
Senior AI Performance & Efficiency Engineer - GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Equity
Comprehensive benefits package
Senior AI Infrastructure Engineer — Scalable GPU Clusters
Senior AI Infrastructure Engineer — Scalable GPU Clusters

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior System Software Engineer - AI Performance and Efficiency Tools
Senior System Software Engineer - AI Performance and Efficiency Tools

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Senior System Software Engineer - AI Performance and Efficiency Tools
Senior System Software Engineer - AI Performance and Efficiency Tools

NVIDIA • United States

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Performance Tools Architect
Senior AI Performance Tools Architect

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Senior AI Performance Profiling Engineer - Equity & Benefits
Senior AI Performance Profiling Engineer - Equity & Benefits

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior HPC AI Cluster Architect — Equity Eligible
Senior HPC AI Cluster Architect — Equity Eligible

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 176,000 - 334,000