Principal High-Performance LLM Training Engineer

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe is looking for a Principal Engineer to enhance performance in AI training and workloads across its hardware and software stack. This role focuses on performance analysis and optimization for LLM workloads running on NVIDIA’s cutting-edge technology.

The successful candidate will need a strong technical background, with over 12 years of experience in relevant fields and a proven record in AI systems. A diverse work environment and robust benefits package await you.

Qualifications

  • 12+ years of relevant work or research experience.
  • Deep hands-on experience analyzing and optimizing large-scale deep learning workloads.
  • Strong expertise in GPU and AI accelerator architecture.
  • Experience with distributed training techniques.

Responsibilities

  • Lead performance analysis and optimization of LLM workloads.
  • Identify bottlenecks across various system elements for optimization.
  • Develop production-quality software and tools to improve performance.
  • Mentor engineers and establish best practices in performance analysis.

Skills

Large-scale AI training systems
GPU performance optimization
Distributed systems
High-performance computing
ML frameworks
Technical leadership

Education

MS, PhD, or equivalent experience in Computer Science, Electrical Engineering, or related field

Tools

Profiling tools
Benchmarking tools
Performance modeling tools

Job description

NVIDIA is seeking a Principal Engineer to drive the performance of large‑scale AI training and post‑training workloads across NVIDIA’s full hardware and software stack. This role sits at the intersection of distributed training, GPU architecture, systems software, deep learning frameworks, and performance engineering. You will analyze and optimize frontier‑scale LLM workloads running on thousands of GPUs, drive improvements across frameworks such as PyTorch, JAX, NeMo, and NeMo RL, and use insights from real workloads to help shape future NVIDIA GPU, system, and software roadmaps.

Responsibilities
  • Lead end-to-end performance analysis and optimization of innovative LLM pre‑training and post‑training workloads on the latest NVIDIA hardware and software platforms.
  • Drive workloads closer to speed‑of‑light performance by identifying and removing bottlenecks across compute, memory, communication, scheduling, parallelism strategy, kernel efficiency, framework overhead, and system‑level scaling.
  • Develop production‑quality software, tools, models, benchmarks, and analysis infrastructure that improve training performance, efficiency, and developer velocity across NVIDIA’s AI software stack.
  • Build and refine performance models, workload characterizations, and simulation methodologies to guide future GPU, networking, system, and software architecture decisions.
  • Serve as a technical authority for AI training performance, partnering closely with teams across GPU architecture, systems, CUDA libraries, compilers, networking, frameworks, product management, and applied AI.
  • Translate workload insights into concrete hardware and software recommendations, and advocate for changes that improve performance and efficiency across the AI ecosystem.
  • Mentor and provide technical leadership to engineers across the organization, helping establish best practices for large‑scale AI performance analysis and optimization.
Qualifications
  • MS, PhD, or equivalent experience in Computer Science, Electrical Engineering, Computer Engineering, or a related field, with 12+ years of relevant work or research experience.
  • Demonstrated principal‑level technical impact in one or more of the following areas: large‑scale AI training systems, GPU performance optimization, distributed systems, high‑performance computing, ML frameworks, compilers/runtimes, or hardware/software co‑design.
  • Deep hands‑on experience analyzing and optimizing performance of large‑scale deep learning workloads, especially transformer‑based models, LLM pre‑training, reinforcement learning, fine‑tuning, or other post‑training workloads.
  • Strong understanding of GPU and AI accelerator architecture from individual accelerators to datacenter‑scale systems.
  • Experience with distributed training techniques such as data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, sequence parallelism, activation checkpointing, mixed precision training, and communication/computation overlap.
  • A strong track record of using profiling, tracing, benchmarking, and performance modeling tools to diagnose complex bottlenecks and drive measurable improvements.
  • Excellent communication and technical leadership skills, with the ability to influence architecture and software decisions across multiple teams without relying on direct authority.

Base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000USD – 431,250USD. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May2,2026.

EEO and Diversity Commitment

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal High-Performance LLM Training Engineer
Principal High-Performance LLM Training Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits
Senior High-Performance LLM Training Engineer
Senior High-Performance LLM Training Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Principal Deep Learning Algorithm Engineer
Principal Deep Learning Algorithm Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Principal Deep Learning Algorithm Engineer
Principal Deep Learning Algorithm Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior AI Training Performance Architect
Senior AI Training Performance Architect

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior High-Performance LLM Training Engineer
Senior High-Performance LLM Training Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000