Senior Performance Engineer – DGX Cloud

Jobtailor

California (MO)

On-site

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior AI performance engineer to analyze end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks. You will design rigorous performance studies and establish baselines, diagnose regressions, and quantify bottlenecks.

Collaborating with DL engineers, platform teams, and GPU architects, you will define benchmarks and communicate findings to influence system design and optimizations across the stack.

Qualifications

  • BS or higher degree in computer science, computer engineering, or a related field (or equivalent experience).
  • 12+ years of experience with strong programming skills in C++ and Python, able to build reliable analysis and automation workflows.
  • Solid foundation in operating systems, computer architecture, and distributed systems.
  • Experience with performance engineering, benchmarking, profiling, and optimization of complex software or systems.
  • Ability to communicate technical findings, prioritize high-impact work, and align across teams.
  • Experience analyzing large-scale AI clusters or distributed training/inference workloads.
  • Experience with CUDA, GPU computing systems, and GPU performance analysis.
  • Hands-on experience with deep learning frameworks such as PyTorch or JAX/XLA.
  • Deep understanding of system-level performance analysis, workload characterization, and optimization.

Responsibilities

  • Analyze end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks.
  • Design and execute rigorous performance studies to establish baselines, diagnose regressions, and quantify bottlenecks.
  • Define performance and efficiency evaluation methodologies, benchmarks, and success metrics for AI workloads.
  • Use profiling, observability, and data analysis to turn performance measurements into actionable optimization plans.
  • Partner with deep learning engineers, platform teams, and GPU architects to validate and deliver performance improvements.
  • Communicate performance findings, trade-offs, and recommendations clearly to influence system and software design decisions.

Skills

C++ programming
Python programming
Performance engineering
Profiling
CUDA
GPU performance analysis
Distributed systems
System-level performance analysis
Data analysis
Automation workflows
PyTorch
JAX/XLA

Education

Bachelor's degree in CS/CE or related field

Tools

PyTorch
JAX/XLA
GPU Computing Systems

Job description

  • Analyze end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks.
  • Design and execute rigorous performance studies to establish baselines, diagnose regressions, and quantify bottlenecks.
  • Define performance and efficiency evaluation methodologies, benchmarks, and success metrics for AI workloads.
  • Use profiling, observability, and data analysis to turn performance measurements into actionable optimization plans.
  • Partner with deep learning engineers, platform teams, and GPU architects to validate and deliver performance improvements.
  • Communicate performance findings, trade-offs, and recommendations clearly to influence system and software design decisions.
Requirements
  • BS or higher degree in computer science, computer engineering, or a related field (or equivalent experience).
  • 12+ years of experience in strong programming skills in C++ and Python, with the ability to build reliable analysis and automation workflows
  • Solid foundation in operating systems, computer architecture, and distributed systems
  • Experience with performance engineering, benchmarking, profiling, and optimization of complex software or systems
  • Ability to communicate technical findings, prioritize high-impact work, and build alignment across teams
  • Experience analyzing large-scale AI clusters or distributed training and inference workloads (Ways to stand out from the crowd).
  • Experience with CUDA, GPU computing systems, and GPU performance analysis.
  • Hands-on experience with deep learning frameworks such as PyTorch or JAX/XLA.
  • Deep understanding of system-level performance analysis, workload characterization, and optimization.
Core Competencies

Expertise in performance engineering and optimization of large-scale AI workloads, with strong programming skills in C++ and Python. Proficient in using deep learning frameworks and GPU computing systems to analyze and enhance system performance.

Highest-signal resume keywords
  • C++ Programming
  • Python Programming
  • Performance Engineering
  • CUDA
  • Deep Learning Frameworks
ATS Optimization Keywords
Hard Skills
  • Performance Analysis
  • Benchmarking
  • Profiling
  • Optimization
  • Distributed Systems
  • Operating Systems
  • Workload Characterization
  • Data Analysis
  • Automation Workflows
  • System-Level Performance Analysis
Soft Skills
  • Communication
  • Collaboration
  • Prioritization
  • Influencing
Industry Keywords
  • AI Workloads
  • Large-Scale Clusters
  • Deep Learning
  • Performance Metrics
Tools & Technologies
  • PyTorch
  • JAX/XLA
  • GPU Computing Systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Senior Software Engineer – Local AI
Senior Software Engineer – Local AI

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Oregon (WI)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Redmond (WA)

On-site
USD 224,000 - 432,000
Equity
Benefits
System Software Engineer – AI
System Software Engineer – AI

Jobtailor • California (MO)

On-site
USD 120,000 - 170,000
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA AI • Eugene (OR)

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
AI and ML Infra Software Engineer, GPU Clusters
AI and ML Infra Software Engineer, GPU Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000