Lead GPU AI Inference Performance Engineer

AMD

San Jose (CA)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading semiconductor company is seeking a Principal AI Performance Engineer in San Jose, CA. The ideal candidate will drive performance optimization on AMD GPUs, lead a technical team, and engage with customers to present findings and recommendations. Candidates should have 7+ years of experience in GPU computing or AI systems, deep knowledge of performance diagnostics, and proficiency in Python and C++. This role requires a strong academic background and the ability to leverage AI tools for performance engineering.

Qualifications

  • 7+ years of experience in software development, especially in GPU computing or AI systems.
  • Hands-on experience with AI serving frameworks and their internals.
  • Ability to understand and optimize GPU kernel performance.

Responsibilities

  • Drive end-to-end performance optimization for AI models.
  • Profile and diagnose performance bottlenecks across multiple systems.
  • Lead technical engagements with clients, presenting findings and recommendations.

Skills

AI fluency
Profiling tools expertise
Python proficiency
C++ proficiency
Technical leadership
Deep understanding of GPU computing
Experience with AI frameworks

Education

Bachelor’s, Master’s, or PhD in relevant field

Tools

TensorRT
CUDA
Triton
vLLM

Job description

A leading semiconductor company is seeking a Principal AI Performance Engineer in San Jose, CA. The ideal candidate will drive performance optimization on AMD GPUs, lead a technical team, and engage with customers to present findings and recommendations. Candidates should have 7+ years of experience in GPU computing or AI systems, deep knowledge of performance diagnostics, and proficiency in Python and C++. This role requires a strong academic background and the ability to leverage AI tools for performance engineering.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Performance Engineer, GPU Systems Lead
Principal AI Performance Engineer, GPU Systems Lead

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
GenAI Inference Optimization Lead — GPU Performance
GenAI Inference Optimization Lead — GPU Performance

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Senior GPU ML Training Performance Engineer
Senior GPU ML Training Performance Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 130,000 - 170,000
Comprehensive health benefits
Inclusive workplace culture
Opportunities for career advancement
Senior GPU Performance Engineer for AI Training
Senior GPU Performance Engineer for AI Training

CareerArc • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits
Lead GPU Performance Engineer for AI Training and Finetuning
Lead GPU Performance Engineer for AI Training and Finetuning

AMD • San Jose (CA)

Hybrid
USD 120,000 - 160,000
Fellow GPU Performance Optimizer for AI Training
Fellow GPU Performance Optimizer for AI Training

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 180,000
Health insurance
Retirement plan
Paid time off
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior GPU Training Performance Engineer
Senior GPU Training Performance Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 120,000 - 160,000
Senior AI Performance Architect — GPU & Network
Senior AI Performance Architect — GPU & Network

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package
GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3