AI Inference Performance & Scale Engineer

AMD

San Jose (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Benefits at a glance

Job summary

AMD in San Jose, CA seeks a performance-obsessed engineer to drive AI inference performance on AMD GPUs. Lead a small team, profile and optimize models end-to-end, and uplift customer engagements with measurable results.

You will diagnose kernel-level bottlenecks, optimize multi-node distributed inference, and upstream efforts to open-source frameworks while collaborating with customers and internal teams.

Qualifications

  • Hands-on experience profiling AI workloads end-to-end.
  • Strong knowledge of GPU kernel performance and occupancy.
  • Experience with AI serving frameworks and custom kernels.
  • Ability to diagnose bottlenecks from user requests to kernels.
  • Excellent English communication and customer-facing leadership.

Responsibilities

  • Drive end-to-end performance optimization across the stack for leading models.
  • Profile, diagnose, and resolve cross-stack bottlenecks from GPU kernels to scheduling.
  • Lead customer-facing technical engagements and present optimization results.
  • Integrate and optimize custom kernels within serving frameworks.
  • Improve distributed inference performance across multi-node deployments.
  • Develop and refine reusable performance optimization methodologies.
  • Leverage AI agents to accelerate daily work and define best practices.

Skills

GPU computing
Python
C++
Performance profiling
AI systems
Linux

Education

Bachelor's degree in CS/CE
Advanced degree preferred

Tools

CUDA
Triton
Gluon
PyTorch
SGLang

Job description

AMD in San Jose, CA seeks a performance-obsessed engineer to drive AI inference performance on AMD GPUs. Lead a small team, profile and optimize models end-to-end, and uplift customer engagements with measurable results.

You will diagnose kernel-level bottlenecks, optimize multi-node distributed inference, and upstream efforts to open-source frameworks while collaborating with customers and internal teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Performance & Reliability Engineer
Senior AI Performance & Reliability Engineer

AMD • San Jose (CA)

Hybrid
USD 180,000 - 260,000
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Fellow AI Performance & Reliability Engineer (Hybrid)
Fellow AI Performance & Reliability Engineer (Hybrid)

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance
AI Systems Engineer: High-Performance ML on Accelerators
AI Systems Engineer: High-Performance ML on Accelerators

AMD • San Jose (CA)

On-site
USD 150,000 - 190,000
Senior GPU/AI Systems Engineer - Performance & ML
Senior GPU/AI Systems Engineer - Performance & ML

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Agentic AI Engineer for Compute & Hardware Optimization
Agentic AI Engineer for Compute & Hardware Optimization

Socket.dev • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Lead Data Center GPU Performance Architect for AI
Lead Data Center GPU Performance Architect for AI

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 250,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits