AI System Performance Engineer (NSP/NPU, AI/ML)

Qualcomm

Bengaluru

On-site

INR 900,000 - 1,700,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Qualcomm Bengaluru is seeking a System Performance Engineer to profile and optimize Snapdragon AI/ML workloads on NPU, CPU, and memory subsystems. You will drive performance analyses using industry benchmarks and collaborate with cross-functional teams to guide optimization efforts across chipsets.

The ideal candidate has 1–5 years of experience in performance profiling on ARM/x86 platforms, with knowledge of Linux/QNX and model optimization techniques.

Qualifications

  • Bachelor's degree in Computer/Electrical Engineering or closely related field.
  • 1–5 years of industry experience in performance engineering on ARM/x86 platforms.
  • Experience with Linux/QNX development and building makefiles is a plus.

Responsibilities

  • Drive performance analysis on silicon using NPU, CPU, and memory benchmarks.
  • Use performance tools to identify bottlenecks in system and IPs.
  • Analyze Perf KPIs of SoC subsystems and correlate with projections.
  • Characterize performance across junction temperatures and optimize for high ambient temps.
  • Analyze and optimize SoC NoC and DDR performance parameters.
  • Collaborate with cross-functional teams to plan performance activities and shape next generation chipsets.

Skills

NPU micro-arch
DL fundamentals
Performance profiling
ARM/x86 platforms
Benchmarks like GeekBench
Perf tools (VTune, perf)
Memory hierarchy
CUDA/ML frameworks

Education

Bachelor's, Computer/EE
Master's, Computer/EE

Tools

QNN
ONNX Runtime
TensorFlow Lite
PyTorch Mobile
Arm/Intel OpenVINO
Qualcomm tools

Job description

Job Area: System Performance Engineer (NSP/NPU, AI/ML)

Job Overview: You will be part of System Performance team that is responsible for profiling and optimizing the System Performance on Snapdragon chipsets. This role will require a strong knowledge of AI/ML performance using NPU cores. The knowledge of CPU, DDR, NoCs will be an added advantage.

Responsibilities
  • Drive Performance analysis on silicon using various System and Cores (i.e. NPU, AI/ML, CPU, Memory) benchmarks like Dhrystone, GeekBench, SPECInt, CNN/GenAI ML networks etc.
  • Use of Performance tools to analyze the load patterns across IPs and identify any performance bottlenecks in system.
  • Analyzing Perf KPIs of SoC subsystems like NPU, CPU, Memory, and corelate performance with projection
  • Evaluate and characterize performance at various junction temperatures and optimize running at high ambient temperatures.
  • Analyze and optimize the System performance parameters of SoC infrastructure like NoC, LP5 DDR, etc.
  • Collaborate with cross-functional global teams to plan and execute performance activities on chipsets as well as make recommendations for next generation chipsets.
Qualifications
  • Minimum Qualifications:
  • 1 to5 years of industry experience in the following:
  • Experience working on any ARM/x86 based platforms, mobile/automotive/Data center operating systems and/or performance profiling tools.
  • Experience in application or driver development in Linux\QNX and ability to create/customize make files with various compiler options is a plus.
  • Must be quick learner and should be able to adapt to new technologies.
  • Must have excellent communication skills.
  • Education Requirements:
  • Required: Bachelor's, Computer Engineering, and/or Electrical Engineering
  • Preferred: Master's, Computer Engineering, and/or Electrical Engineering
Required Skills
  • A deep understanding of NPU, CPU and DDR architecture internals like
  • NPU: Understanding of NPU micro-architecture, including its specialized processing elements for operations like matrix multiplications and convolutions.
  • A strong grasp of deep learning fundamentals, including neural network architectures like Convolutional Neural Networks (CNNs), transformers, GenAI etc.
  • Understanding the distinction between the compute-intensive training of a model (often done on GPUs) and running inference (making predictions) on NPUs.
  • Knowledge of Model optimization and compression techniques (Quantization, Pruning, Distillation etc.)
  • CPU caches (L1, L2, L3), instruction pipelines, branch prediction, memory hierarchy (including register, cache, and main memory) and multi-core/multi-threaded processing. You need to understand how memory access, instruction dependencies, and contention for shared resources can impact performance.
  • DDR: Understanding of JEDEC specifications LPDDR4, LPDDR5, LPDDR6, including command and timing parameters (e.g., 𝑡𝐶𝐿, 𝑡𝑅𝐶𝐷, tRP), memory organization (rows, columns, banks), and basic view of training and initialization sequences, how a memory controller works and its specific features like command queue, port arbitration, and various control schemes.
  • Familiarity with frameworks like TensorFlow Lite, PyTorch Mobile, and ONNX Runtime for preparing models for deployment. Experience with NPU-specific compilers, such as those from Qualcomm (QNN), Arm (Vela), Intel (OpenVINO), to optimize and orchestrate AI workloads.
  • Evaluating hardware and software to determine the best fit for specific AI tasks and measure performance metrics like TOPS/W (trillions of operations per second perwatt).
  • Using tools like the Windows Performance Toolkit and Qualcomm's Snapdragon Profiler to analyze and optimize NPU and system-level performance A core understanding of how to use parallel processing architectures for efficient AI computations.
  • Expertise in how operating systems manage processes, threads, memory, MMU, and interrupt handling. This knowledge is crucial to understanding software for the kernel scheduler and system-level bottlenecks.
  • Good understanding of Benchmarks CPU (like GeekBench, SpecInt, CoreMark etc) and DDR (like lat_mem, stream, bw_mem etc.) and how they exercise the underlying CPU/GPU/DDR architecture.
  • Experience with a variety of performance monitoring tools like Intel VTune, Linux perf, and Utilities like top, vmstat, iostat, and netstat to monitor system resources like CPU, memory, and I/O. Experience with software tools to monitor system hotspots, command bus utilization, and identify memory traffic patterns is critical. This includes validating that the traffic generated by software is as expected.
  • Good understanding of memory allocation policies, prefetching, and caching to minimize latency and maximize bandwidth. Understanding how an application accesses memory is vital. Skills in profiling code to analyze memory access patterns and then optimizing the code for better data locality.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

System Performance Engineer (NSP/NPU, AI/ML)
System Performance Engineer (NSP/NPU, AI/ML)

QUALCOMM, Inc. • Bengaluru

On-site
INR 900,000 - 1,500,000
AI Infrastructure System Performance Architect (IRT815ST RM 4407)
AI Infrastructure System Performance Architect (IRT815ST RM 4407)

Source-Right • Hyderabad

On-site
INR 4,200,000 - 7,000,000
Architect - System Performance Verification and Analysis
Architect - System Performance Verification and Analysis

NVIDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead II - Embedded Software
Lead II - Embedded Software

UST • Bengaluru

On-site
INR 1,000,000 - 1,500,000
NPU/AI Processor Architecture (RTL)/Arithmetic Design Engineer - Sr Staff
NPU/AI Processor Architecture (RTL)/Arithmetic Design Engineer - Sr Staff

Qualcomm • Bengaluru

On-site
INR 1,200,000 - 1,800,000
AI Model System Software Performance Optimization Engineer - Senior Engineer
AI Model System Software Performance Optimization Engineer - Senior Engineer

QUALCOMM, Inc. • Bengaluru

On-site
INR 1,500,000 - 2,600,000
Edge AI Engineer
Edge AI Engineer

Vedya Labs • Hyderabad, Bengaluru

On-site
INR 900,000 - 1,500,000
AI Accelerator Chip Architect
AI Accelerator Chip Architect

Capgemini Engineering • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Performance Benchmark– Server Platform - Engineer to Senior Staff Engineer
Performance Benchmark– Server Platform - Engineer to Senior Staff Engineer

Qualcomm • Bengaluru

On-site
INR 800,000 - 1,200,000
Server Performance Architect - Hardware
Server Performance Architect - Hardware

NVIDIA • Hyderabad

On-site
INR 4,000,000 - 7,000,000