Senior GPU Inference Performance Engineer

Advanced Micro Devices

Santa Clara (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Advanced Micro Devices is seeking a Senior GPU Inference Performance Engineer to own end-to-end profiling of GPU-accelerated AI inference workloads. You will analyze workloads across AMD Instinct and NVIDIA GPUs, profile AI serving frameworks, and explain performance gaps with hardware and software evidence.

You will collaborate across teams on multi-server inference networking and Kubernetes-related performance concerns.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred.

Responsibilities

  • Full-stack GPU profiling across AMD Instinct and NVIDIA GPUs.
  • Profile and optimize AI serving frameworks (e.g., vLLM, SGLang) and quantify throughput/latency implications.
  • Design and run head-to-head benchmarks to explain performance differences with hardware/software evidence.
  • Profile multi-server inference networking and distributed topologies across clusters.
  • Identify and mitigate latency/jitter due to GPU operator, container, and Kubernetes stack.

Skills

GPU performance engineering
HPC performance analysis
Python
C/C++
Kubernetes GPU scheduling
LLM/ML serving profiling
Technical communication

Education

Bachelor's degree in Computer Science/Engineering or related field

Tools

ROCm / rocProfiler / Omniperf
CUDA / Nsight Systems / Nsight Compute
RCCL / NCCL
RDMA/RoCE tracing tools
Memory Fabric profiling tools

Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.Together, we advance your career.

THE ROLE:

We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and explain performance across the full stack, from GPU silicon through the software runtime, and drive competitive positioning against other accelerator vendors. This role sits at the intersection of hardware, systems software, and AI serving frameworks, and requires someone who can go deep on a trace and present findings to product and executive stakeholders.

THE PERSON:

A hands‑on performance engineer who is equally comfortable reading a GPU trace and briefing executives. You are curious, evidence‑driven, rigorous and you don't stop at "X is faster," you explain why, rooted in hardware and software evidence. You collaborate across hardware, systems software, and AI serving framework teams, communicate clearly in written reports and presentations, and thrive at the intersection of silicon, systems, and AI.

KEY RESPONSIBILITIES:
  • Full‑stack GPU profiling: Instrument and analyze inference workloads across AMD Instinct (ROCm, rocProfiler, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs. Identify bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and PCIe/Infinity Fabric data movement.
  • AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging serving runtimes. Understand KV‑cache management, continuous batching, PagedAttention, speculative decoding, and quantization (FP8, MXFP4, INT4) effects on throughput and latency.
  • Competitive performance analysis: Design and execute head‑to‑head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data‑backed explanations of why performance differs — attributing gaps to specific hardware features (HBM bandwidth, compute density, interconnect topology), software maturity (kernel libraries, operator fusion, graph compilation), or configuration differences.
  • Multi‑server inference networking: Profile and optimize distributed inference topologies including prefill‑decode (PD) disaggregation, pipeline parallelism, and tensor parallelism across multi‑node clusters. Analyze network‑level bottlenecks using RDMA/RoCE traces, NCCL/RCCL collective profiling, and NIC‑level counters (Pensando, ConnectX). Quantify the impact of network latency, bandwidth, and congestion on end‑to‑end inference SLAs.
  • GPU operator and Kubernetes stack: Profile the overhead introduced by GPU operators, device plugins, container runtimes (Docker, containerd), and Kubernetes scheduling on inference latency. Identify and resolve jitter, cold‑start, and resource contention issues in production serving environments.
  • Tooling and automation: Build reproducible benchmarking harnesses, profiling scripts, and performance regression dashboards. Automate trace collection and analysis to support continuous performance validation across driver, firmware, and framework updates.
PREFERRED EXPERIENCE:
  • Background in GPU performance engineering, HPC, or systems performance analysis
  • Hands‑on proficiency with either AMD (ROCm, rocProfiler, Omniperf/Omnitrace) or NVIDIA (CUDA, Nsight Systems/Compute, NCU) profiling toolchains, with deep understanding of GPU architecture: warp/wavefront execution, memory hierarchy (registers → LDS/shared → L2 → HBM), occupancy, and instruction‑level parallelism
  • Experience profiling vLLM, SGLang, or equivalent LLM serving frameworks, including quantization workflows (FP8, MXFP4, INT4, AWQ, GPTQ) and their performance implications
  • Experience with multi‑GPU and multi‑node inference — tensor parallelism, pipeline parallelism, or PD disaggregation over RDMA/RoCE — including RCCL/NCCL profiling and network tools (perftest, ib_write_bw, tcpdump, Memory Fabric counters)
  • Demonstrated ability to explain performance differences in written reports or presentations — not just "X is faster" but why, rooted in hardware and software evidence
  • Strong Python and C/C++ skills; comfort reading GPU kernel code (HIP/CUDA)
  • Experience with Kubernetes GPU scheduling, MIG, and GPU operator performance, or contributions to open‑source inference or profiling projects
ACADEMIC CREDENTIALS:
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred; advanced degree desired

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s "Responsible AI Policy" is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Senior GPU Software Performance Engineer — Post‐Training
Senior GPU Software Performance Engineer — Post‐Training

Advanced Micro Devices • San Jose (CA)

On-site
USD 130,000 - 170,000
Comprehensive health benefits
Inclusive workplace culture
Opportunities for career advancement
Senior GPU Software Performance Engineer – Post-Training
Senior GPU Software Performance Engineer – Post-Training

AMD • San Jose (CA)

On-site
Principal / Senior GPU SW Performance Engineer — Post‑Training
Principal / Senior GPU SW Performance Engineer — Post‑Training

AMD • San Jose (CA)

On-site
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Benefits at a glance
Data Center GPU Performance Attainment Lead
Data Center GPU Performance Attainment Lead

Advanced Micro Devices • Austin (TX)

Hybrid
USD 90,000 - 120,000
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Competitive salary
Health benefits
Career advancement opportunities
Principal Software Development Eng. - AI Performance
Principal Software Development Eng. - AI Performance

Advanced Micro Devices • San Jose (CA)

On-site
USD 190,000 - 230,000
Principal GenAI Inference Optimization Engineer
Principal GenAI Inference Optimization Engineer

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Fellow GPU Performance Optimization Engineer
Fellow GPU Performance Optimization Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive benefits