AI & HPC GPU Compute Performance Engineer

engineeringjobs.net, Inc.

San Jose (CA)

On-site

USD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Matching
Paid Parental Leave
Short-Term Disability Coverage
Long-Term Disability Coverage
Life Insurance
Restricted Stock Units
Paid Holidays
Floating Holiday
Birthday Leave
Year-End Holiday Shutdown
Personal Wellness Days
Paid Vacation
Sick Leave
Family Emergency Leave
Paid Volunteer Days
Annual Bonus

Job summary

engineeringjobs.net, Inc. is seeking an experienced engineer to run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems.

You will use telemetry and profiling tools to identify bottlenecks across compute, memory, interconnect, and storage, and document procedures and results. You will build automation for performance testing, regression detection, and failure triage, and collaborate with hardware and software teams to resolve issues.

Qualifications

  • Requires degree with related experience or equivalent work history.
  • Linux systems debugging, testing, or tuning experience.
  • Experience evaluating AI, ML, or HPC workloads.
  • 3+ years with a programming or scripting language.

Responsibilities

  • Run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems.
  • Use telemetry and profiling tools to identify bottlenecks across compute, memory, interconnect, and storage.
  • Build automation for performance testing, regression detection, and failure triage.
  • Collaborate with hardware and software teams to resolve issues.
  • Document procedures and results.

Skills

Linux Systems Debugging
AI Workload Performance Analysis
HPC Workloads
Python
C
C++
Go
Bash
GPU Computing
CUDA
ROCm
NCCL
RCCL
Performance Profiling
Kubernetes
Slurm

Education

Bachelor's degree
Master's degree
PhD with relevant research
Equivalent work experience

Tools

Kubernetes
Slurm

Job description

engineeringjobs.net, Inc. is seeking an experienced engineer to run and analyze AI training, inference, and HPC workloads on GPU-accelerated systems.

You will use telemetry and profiling tools to identify bottlenecks across compute, memory, interconnect, and storage, and document procedures and results. You will build automation for performance testing, regression detection, and failure triage, and collaborate with hardware and software teams to resolve issues.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Compute Performance Engineer: GPU & HPC
AI Compute Performance Engineer: GPU & HPC

Cisco Systems, Inc • Milpitas (CA)

On-site
USD 155,000 - 223,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+3
AI Compute Performance Engineer
AI Compute Performance Engineer

020 Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 155,000 - 223,000
AI Compute Performance Engineer
AI Compute Performance Engineer

Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 155,000 - 223,000
Medical, dental, vision insurance
401(k) with Cisco matching
Paid parental leave
+2
AI Performance Engineer – HPC, ARM & Distributed Inference
AI Performance Engineer – HPC, ARM & Distributed Inference

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000
AI Inference & HPC Engineer – Performance & APIs
AI Inference & HPC Engineer – Performance & APIs

Topaz Labs • Emeryville (CA)

On-site
USD 90,000 - 150,000
Full medical/dental/vision coverage
15 days PTO
5 personal days + holidays
+2
Software Engineer, Compute Performance
Software Engineer, Compute Performance

engineeringjobs.net, Inc. • San Jose (CA)

On-site
USD 150,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+16
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior AI/HPC GPU Engineer - Performance & Systems Expert
Senior AI/HPC GPU Engineer - Performance & Systems Expert

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Senior GPU HPC Infra Engineer for AI Training
Senior GPU HPC Infra Engineer for AI Training

Showcify • United States

On-site
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
RRSP/401K
+4