Lead HPC Software Optimization Engineer - C++

AMD

India

On-site

INR 450,000 - 750,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits at a glance

Job summary

AMD is seeking a Senior Software Developer to lead GPU kernel optimization for AI workloads. You will architect/implement optimized compute kernels, design multi-GPU/multi-node scaling, profile systems for max hardware utilization, and build benchmarking infrastructure.

You’ll guide agile teams, work with PyTorch, vLLM, Cutlass, Kokkos, and profiling tools, and ship production software through upstreaming or rollouts. Strong C++ and CUDA development are required.

Qualifications

  • Bachelor’s or Master’s degree in Computer Engineering, Electrical Engineering, Computer Science, or related field.

Responsibilities

  • GPU Kernel Optimization: Develop and optimize GPU kernels for AI workloads.
  • Multi-GPU and Multi-Node Scaling: Architect distributed training/inference strategies across GPUs/nodes.
  • Performance Profiling: Use tools to identify bottlenecks and optimize hardware utilization.
  • Parallel Computing: Implement multi-threaded techniques for scalable execution on GPUs.
  • Benchmarking & Testing: Build infrastructure to assess performance and reliability.
  • Documentation & Best Practices: Produce technical docs and reusable components.

Skills

GPU kernel development
CUDA
HIP
C++20
Distributed computing
Performance profiling
Python
Multi-GPU scaling
GEMMs

Education

Bachelor’s or Master’s degree in Computer Engineering/CS/EE

Tools

NCCL
MPI
ROCm/CUDA runtimes
CUDA Toolkit
PyTorch
Kokkos
CUTLASS

Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.

THE TEAM:

Join AMD’s high-impact team at the heart of innovation in AI, ML, and high-performance computing (HPC). We’re a collaborative group of software architects and GPU engineers focused on pushing the boundaries of AI model performance across distributed, GPU-accelerated platforms. Our work drives the next generation of AMD’s AI software stack, enabling large-scale machine learning training and inference workloads in data centers and enterprise environment

THE ROLE:

AMD is hiring a Senior Software Developer for its AI, ML, and high-performance computing team to lead GPU kernel optimization and distributed software for large-scale AI workloads. In this technical leadership role, you'll architect and implement optimized compute kernels, design multi-GPU/multi-node scaling strategies, profile systems to maximize hardware utilization, build benchmarking infrastructure, and guide agile teams across the product lifecycle. The ideal candidate is a deep systems thinker fluent in GPU architecture, parallel computing, and AI model execution, comfortable both writing performance‑critical code and driving software architecture decisions. Required expertise includes GPU kernel optimization in C++ (17/20), hands‑on CUDA and low‑level GPU programming, distributed AI computing (multi‑GPU, NCCL, MPI), and familiarity with frameworks such as PyTorch, vLLM, Cutlass, and Kokkos, plus strong performance tuning with profiling tools (Nsight, VTune, Perf) and Python automation. You should also bring proven software leadership experience—defining roadmaps with stakeholders, interfacing with executives, and shipping production software through open‑source upstreaming or commercial rollouts.

THE PERSON:

We’re looking for a highly skilled, deep systems thinker who thrives in complex problem domains involving parallel computing, GPU architecture, and AI model execution. You are confident leading software architecture decisions and know how to translate business goals into robust, optimized software solutions. You’re just as comfortable writing performance‑critical code as you are guiding agile development teams across product lifecycles. Ideal candidates have a strong balance of low‑level programming, distributed systems knowledge, and leadership experience—paired with a passion for AI performance at scale.

KEY RESPONSIBILITIES:
  • GPU Kernel Optimization: Develop and optimize GPU kernels to accelerate inference and training of large machine learning models while ensuring numerical accuracy and runtime efficiency.
  • Multi‑GPU and Multi‑Node Scaling: Architect and implement strategies for distributed training/inference across multi‑GPU/multi‑node environments using model/data parallelism techniques.
  • Performance Profiling: Identify bottlenecks and performance limitations using profiling tools; propose and implement optimizations to improve hardware utilization.
  • Parallel Computing: Design and implement multi-threaded and synchronized compute techniques for scalable execution on modern GPU architectures.
  • Benchmarking & Testing: Build robust benchmarking and validation infrastructure to assess performance, reliability, and scalability of deployed software.
  • Documentation & Best Practices: Produce technical documentation and share architectural patterns, code optimization tips, and reusable components.
PREFERRED EXPERIENCE:
  • GPU kernel development (HIP, CUDA C/C++, PTX, GPU Assembly)
  • GPU kernel optimization down to assembly level
  • GPU hardware architectures (AMD, nVidia, Intel, …)
  • ML/HPC related parallel algorithm design, e.g., GEMMs, element‑wise, attention, reductions
  • ROCm/CUDA Software Stacks (Runtimes, Compilers, Libraries)
  • Scripting knowledge (ex. Python)
  • Advanced C++ software development (including meta‑programming, C++20 features)
  • Software development, analyzing, and debugging of complex algorithms.
  • Reading, understanding, and changing advanced C++
  • Reading, understanding, and changing complex assembly
  • Advanced knowledge of software development processes.
  • Excellent working knowledge of GPU/CPU architectures.
  • Distributed computing and multi‑GPU environments.
  • Advanced performance profiling and optimization tools.
  • C++ Performance optimization
  • Optimizing GPU kernels in C++20.
  • Strong experience in low‑level GPU kernel optimization.
  • Optimization of GPU assembly
  • Practical usage of the LLVM compiler flow and tools
  • Proficiency in HIP / CUDA and GPU programming.
  • GPU performance bottleneck analysis
  • GPU Power Optimization analysis
  • Understanding Neural Network models data flow and operations
  • Working with complex frameworks/libraries such as (PyTorch, vLLM, CUTLASS, Kokkos, etc.)
  • Working on complex algorithms at operator level (variances of attention algorithms,MoE, quantization/scalingetc.)
  • OS Kernel Debug and Kernel optimization
  • GPU API programming
  • CPU/ASIC software development
ACADEMIC CREDENTIALS
  • Bachelor’s or Master’s degree in Computer Engineering, Electrical Engineering, Computer Science, or a related technical field.
  • Advanced degrees or published work in HPC, GPU computing, or AI systems is a plus.

#LI-NR1

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead HPC Software Optimization Engineer - C++
Lead HPC Software Optimization Engineer - C++

Advanced Micro Devices • Hyderabad

On-site
INR 3,000,000 - 6,000,000
AMD benefits at a glance
Lead HPC Software Optimization Engineer - C++
Lead HPC Software Optimization Engineer - C++

AMD • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Benefits described by AMD
Staff/Principal GPU Kernel Optimization Engineer
Staff/Principal GPU Kernel Optimization Engineer

AMD • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Lead Open Source AI/ML Solutions Engineer
Lead Open Source AI/ML Solutions Engineer

AMD • Bengaluru

On-site
INR 4,000,000 - 5,200,000
DC-GPU Performance Modeling Engineer
DC-GPU Performance Modeling Engineer

Advanced Micro Devices • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Comprehensive health benefits
Inclusive work culture
Opportunities for professional development
Software Development Engineer
Software Development Engineer

Advanced Micro Devices • Bengaluru

On-site
INR 2,200,000 - 3,200,000
DC-GPU Performance Modeling Engineer
DC-GPU Performance Modeling Engineer

AMD • Hyderabad

On-site
INR 1,000,000 - 1,500,000
GPU Performance Modeling & Optimization Engineer
GPU Performance Modeling & Optimization Engineer

AMD • Bengaluru Urban

On-site
INR 4,000,000 - 7,000,000
AMD benefits
Technical Lead Computer Vision Software Development
Technical Lead Computer Vision Software Development

AMD • India

On-site
INR 2,000,000 - 3,000,000
DC-GPU Performance Modeling Engineer
DC-GPU Performance Modeling Engineer

AMD • Hyderabad

On-site
INR 3,500,000 - 7,000,000
AMD benefits at a glance