Lead HPC Software Optimization Engineer - C++

Advanced Micro Devices

Hyderabad

On-site

INR 3,000,000 - 6,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits at a glance

Job summary

Advanced Micro Devices (AMD) is seeking a Senior Software Developer to lead GPU kernel optimization and distributed software for large‑scale AI workloads. You will architect and implement compute kernels, design multi‑GPU/multi‑node scaling, and guide agile teams through the product lifecycle.

The role requires deep knowledge of C++ (20), CUDA and ROCm, strong performance profiling, and hands‑on experience with PyTorch and related frameworks.

Qualifications

  • Bachelor's or Master's in CS/Engineering; advanced degrees welcome.
  • Strong knowledge of GPU architectures and HPC concepts.
  • Experience with multi-GPU scaling, distributed training, and inference.

Responsibilities

  • Develop and optimize GPU kernels for AI workloads.
  • Design multi-GPU/multi-node distribution strategies.
  • Profile performance and improve hardware utilization.
  • Lead software architecture decisions and mentor teams.
  • Build benchmarking and validation tools for AI systems.
  • Document patterns and best practices for reuse.

Skills

GPU kernel development
CUDA
C++20
Distributed computing
Performance profiling
Leadership experience
Python

Education

Bachelor's or Master's in Computer Science/Engineering
Advanced degrees a plus

Tools

Nsight
VTune
MPI

Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next‑generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.

THE TEAM:

Join AMD’s high-impact team at the heart of innovation in AI, ML, and high-performance computing (HPC). We’re a collaborative group of software architects and GPU engineers focused on pushing the boundaries of AI model performance across distributed, GPU‑accelerated platforms. Our work drives the next generation of AMD’s AI software stack, enabling large‑scale machine learning training and inference workloads in data centers and enterprise environments.

THE ROLE:

AMD is hiring a Senior Software Developer for its AI, ML, and high‑performance computing team to lead GPU kernel optimization and distributed software for large‑scale AI workloads. In this technical leadership role, you'll architect and implement optimized compute kernels, design multi‑GPU/multi‑node scaling strategies, profile systems to maximize hardware utilization, build benchmarking infrastructure, and guide agile teams across the product lifecycle. The ideal candidate is a deep systems thinker fluent in GPU architecture, parallel computing, and AI model execution, comfortable both writing performance‑critical code and driving software architecture decisions. Required expertise includes GPU kernel optimization in C++ (17/20), hands‑on CUDA and low‑level GPU programming, distributed AI computing (multi‑GPU, NCCL, MPI), and familiarity with frameworks such as PyTorch, vLLM, Cutlass, and Kokkos, plus strong performance tuning with profiling tools (Nsight, VTune, Perf) and Python automation. You should also bring proven software leadership experience—defining roadmaps with stakeholders, interfacing with executives, and shipping production software through open‑source upstreaming or commercial rollouts.

THE PERSON:

We’re looking for a highly skilled, deep systems thinker who thrives in complex problem domains involving parallel computing, GPU architecture, and AI model execution. You are confident leading software architecture decisions and know how to translate business goals into robust, optimized software solutions. You’re just as comfortable writing performance‑critical code as you are guiding agile development teams across product lifecycles. Ideal candidates have a strong balance of low‑level programming, distributed systems knowledge, and leadership experience—paired with a passion for AI performance at scale.

KEY RESPONSIBILITIES:
  • GPU Kernel Optimization: Develop and optimize GPU kernels to accelerate inference and training of large machine learning models while ensuring numerical accuracy and runtime efficiency.
  • Multi‑GPU and Multi‑Node Scaling: Architect and implement strategies for distributed training/inference across multi‑GPU/multi‑node environments using model/data parallelism techniques.
  • Performance Profiling: Identify bottlenecks and performance limitations using profiling tools; propose and implement optimizations to improve hardware utilization.
  • Parallel Computing: Design and implement multi‑threaded and synchronized compute techniques for scalable execution on modern GPU architectures.
  • Benchmarking & Testing: Build robust benchmarking and validation infrastructure to assess performance, reliability, and scalability of deployed software.
  • Documentation & Best Practices: Produce technical documentation and share architectural patterns, code optimization tips, and reusable components.
PREFERRED EXPERIENCE:
  • GPU kernel development (HIP, CUDA C/C++, PTX, GPU Assembly)
  • GPU kernel optimization down to assembly level
  • GPU hardware architectures (AMD, nVidia, Intel)
  • ML/HPC related parallel algorithm design, e.g., GEMMs, element‑wise, attention, reductions
  • ROCm/CUDA Software Stacks (Runtimes, Compilers, Libraries)
  • Scripting knowledge (ex. Python)
  • Advanced C++ software development (including meta‑programming, C++20 features)
  • Software development, analyzing, and debugging of complex algorithms.
  • Reading, understanding, and changing advanced C++
  • Reading, understanding, and changing complex assembly
  • Advanced knowledge of software development processes.
  • Excellent working knowledge of GPU/CPU architectures.
  • Distributed computing and multi‑GPU environments.
  • Advanced performance profiling and optimization tools.
  • C++ Performance optimization
  • Optimizing GPU kernels in C++20.
  • Strong experience in low‑level GPU kernel optimization.
  • Optimization of GPU assembly
  • Practical usage of the LLVM compiler flow and tools
  • Proficiency in HIP / CUDA and GPU programming.
  • GPU performance bottleneck analysis
  • GPU Power Optimization analysis
  • Understanding Neural Network models data flow and operations
  • Working with complex frameworks/libraries such as (PyTorch, vLLM, CUTLASS, Kokkos)
  • Working on complex algorithms at operator level (variances of attention algorithms, MoE, quantization/scaling etc.)
  • OS Kernel Debug and Kernel optimization
  • GPU API programming
  • CPU/ASIC software development
ACADEMIC CREDENTIALS
  • Bachelor’s or Master’s degree in Computer Engineering, Electrical Engineering, Computer Science, or a related technical field.
  • Advanced degrees or published work in HPC, GPU computing, or AI systems is a plus.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead HPC Software Optimization Engineer - C++
Lead HPC Software Optimization Engineer - C++

AMD • India

On-site
INR 450,000 - 750,000
AMD benefits at a glance
Lead HPC Software Optimization Engineer - C++
Lead HPC Software Optimization Engineer - C++

AMD • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Benefits described by AMD
Staff/Principal GPU Kernel Optimization Engineer
Staff/Principal GPU Kernel Optimization Engineer

AMD • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Lead Open Source AI/ML Solutions Engineer
Lead Open Source AI/ML Solutions Engineer

AMD • Bengaluru

On-site
INR 4,000,000 - 5,200,000
Software Development Engineer
Software Development Engineer

Advanced Micro Devices • Bengaluru

On-site
INR 2,200,000 - 3,200,000
DC-GPU Performance Modeling Engineer
DC-GPU Performance Modeling Engineer

Advanced Micro Devices • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Comprehensive health benefits
Inclusive work culture
Opportunities for professional development
DC-GPU Performance Modeling Engineer
DC-GPU Performance Modeling Engineer

AMD • Hyderabad

On-site
INR 1,000,000 - 1,500,000
GPU Performance Modeling & Optimization Engineer
GPU Performance Modeling & Optimization Engineer

AMD • Bengaluru Urban

On-site
INR 4,000,000 - 7,000,000
AMD benefits
Lead Performance and Optimization Engineer
Lead Performance and Optimization Engineer

AMD • Bengaluru Urban

On-site
INR 900,000 - 1,500,000
Lead Performance Modeling Engineer – GPU SoC
Lead Performance Modeling Engineer – GPU SoC

Advanced Micro Devices • Bengaluru

On-site
INR 1,600,000 - 2,200,000
Comprehensive benefits package