Software Engineer, Compute Performance

engineeringjobs.net, Inc.

San Jose (CA)

On-site

USD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Matching
Paid Parental Leave
Short-Term Disability Coverage
Long-Term Disability Coverage
Life Insurance
Restricted Stock Units
Paid Holidays
Floating Holiday
Birthday Leave
Year-End Holiday Shutdown
Personal Wellness Days
Paid Vacation
Sick Leave
Family Emergency Leave
Paid Volunteer Days
Annual Bonus

Job summary

engineeringjobs.net, Inc. is seeking an experienced engineer to run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems.

You will use telemetry and profiling tools to identify bottlenecks across compute, memory, interconnect, and storage, and document procedures and results. You will build automation for performance testing, regression detection, and failure triage, and collaborate with hardware and software teams to resolve issues.

Qualifications

  • Requires degree with related experience or equivalent work history.
  • Linux systems debugging, testing, or tuning experience.
  • Experience evaluating AI, ML, or HPC workloads.
  • 3+ years with a programming or scripting language.

Responsibilities

  • Run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems.
  • Use telemetry and profiling tools to identify bottlenecks across compute, memory, interconnect, and storage.
  • Build automation for performance testing, regression detection, and failure triage.
  • Collaborate with hardware and software teams to resolve issues.
  • Document procedures and results.

Skills

Linux Systems Debugging
AI Workload Performance Analysis
HPC Workloads
Python
C
C++
Go
Bash
GPU Computing
CUDA
ROCm
NCCL
RCCL
Performance Profiling
Kubernetes
Slurm

Education

Bachelor's degree
Master's degree
PhD with relevant research
Equivalent work experience

Tools

Kubernetes
Slurm

Job description

Run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems, using telemetry and profiling tools to identify compute, memory, interconnect, storage, and scaling bottlenecks. Build automation for performance testing, regression detection, and failure triage, collaborate with hardware and software teams to resolve issues, and document procedures and results.

Requirements: Requires a bachelor’s degree and at least five years of related experience, a master’s degree and at least three years, a PhD with relevant research experience, or equivalent work experience. Candidates need Linux systems debugging, testing, or tuning experience; experience evaluating AI, machine learning, or HPC workloads; and at least three years with a programming or scripting language, while GPU technologies, profiling tools, distributed workload environments, and GPU system architecture are preferred.

Key Skills: Linux Systems Debugging, AI Workload Performance Analysis, HPC Workloads, Python, C, C++, Go, Bash, GPU Computing, CUDA, ROCm, NCCL, RCCL, Performance Profiling, Kubernetes, Slurm

Benefits: Medical Insurance, Dental Insurance, Vision Insurance, 401(k) Matching, Paid Parental Leave, Short-Term Disability Coverage, Long-Term Disability Coverage, Life Insurance, Restricted Stock Units, Paid Holidays, Floating Holiday, Birthday Leave, Year-End Holiday Shutdown, Personal Wellness Days, Paid Vacation, Sick Leave, Family Emergency Leave, Paid Volunteer Days, Annual Bonus

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Equity
Comprehensive benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

On-site
USD 200,000 - 300,000
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Austin (TX)

On-site
USD 272,000 - 431,250
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Redmond (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Oregon (WI)

On-site
USD 224,000 - 431,250
Equity
Benefits
HPC Performance and Validation Engineer
HPC Performance and Validation Engineer

Gosnaphop • Dallas (TX)

On-site
USD 180,000 - 260,000
100% paid medical, dental, vision
401(k)
25 days PTO
+3
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA • Washington

On-site
USD 224,000 - 431,250
Equity
Benefits
Kernel Engineer
Kernel Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000