AI Systems Engineer: High-Performance ML on Accelerators

Advanced Micro Devices

San Jose (CA)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD in San Jose, CA is seeking an AI Systems Engineer to develop and optimize ML workloads on next-generation AMD AI accelerators. You will work at the hardware-software boundary, designing high-performance ML operator kernels, optimizing dataflow pipelines, and enabling industry-leading AI inference across AMD NPU and GPU platforms.

You will collaborate closely with compiler, runtime, silicon, and architecture teams from kernel development to silicon bring-up, translating technical insights

Qualifications

  • Strong software development experience using C/C++ and Python.
  • Experience with parallel programming and performance optimization.
  • Knowledge of ML inference workloads and common operators (GEMM, convolution, attention, softmax).
  • Familiarity with AI frameworks and runtimes such as PyTorch, ONNX Runtime, ROCm.
  • Understanding of computer architecture, memory hierarchies, cache behavior, and accelerator programming models.
  • Experience developing software for GPUs/NPUs/AI accelerators or HPC platforms.
  • Exposure to MLIR/LLVM and compiler technologies.
  • Familiarity with quantization techniques (INT8, FP8, FP16, BF16).

Responsibilities

  • Develop and optimize ML operator kernels and dataflow libraries for AMD AI accelerators.
  • Profile workloads, identify performance bottlenecks, and drive optimizations.
  • Enable and validate ML models within production inference frameworks and runtimes.
  • Collaborate with compiler, runtime, architecture, and silicon teams to deliver high-performance AI solutions.
  • Debug and resolve issues spanning kernel implementation, runtime integration, model accuracy, and hardware bring-up.
  • Contribute to hardware-software co-design by evaluating architectural tradeoffs.
  • Drive innovation in performance methodologies, benchmarking, tooling, and AI system optimization.

Skills

C/C++
Python
Parallel programming
Performance optimization
GPU/accelerator programming
Linux development
Compiler knowledge
PyTorch

Education

Master's or PhD in CS/EE/CE

Tools

PyTorch
ONNX Runtime
ROCm
MLIR
LLVM

Job description

AMD in San Jose, CA is seeking an AI Systems Engineer to develop and optimize ML workloads on next-generation AMD AI accelerators. You will work at the hardware-software boundary, designing high-performance ML operator kernels, optimizing dataflow pipelines, and enabling industry-leading AI inference across AMD NPU and GPU platforms.

You will collaborate closely with compiler, runtime, silicon, and architecture teams from kernel development to silicon bring-up, translating technical insights

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Engineer: High-Performance ML on Accelerators
AI Systems Engineer: High-Performance ML on Accelerators

AMD • San Jose (CA)

On-site
USD 150,000 - 190,000
AI Systems Engineer: ML Kernels & HPC Acceleration
AI Systems Engineer: ML Kernels & HPC Acceleration

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Systems Design Engineer (AI, Software)
Systems Design Engineer (AI, Software)

AMD • San Jose (CA)

On-site
USD 150,000 - 190,000
Systems Design Engineer (AI, Software)
Systems Design Engineer (AI, Software)

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Systems Design Engineer (AI, Software)
Systems Design Engineer (AI, Software)

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 190,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI/ML Compiler Engineer, High-Performance GPUs
AI/ML Compiler Engineer, High-Performance GPUs

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 260,000
Principal MLIR Compiler Architect for AI Acceleration
Principal MLIR Compiler Architect for AI Acceleration

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 240,000
Applied AI Engineer: AI for Hardware & Software
Applied AI Engineer: AI for Hardware & Software

AMD • California (MO)

On-site
USD 180,000 - 240,000
Senior AI Performance & Reliability Engineer
Senior AI Performance & Reliability Engineer

AMD • San Jose (CA)

Hybrid
USD 180,000 - 260,000