AI Accelerator Software Engineer - Graph & Compiler

Ampere

Santa Clara (CA)

On-site

USD 159,000 - 239,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401K retirement plan
Unlimited flextime

Job summary

Ampere is seeking a Software Engineer for its AI Accelerator team to optimize deep learning graphs and maximize performance on Ampere hardware. You will work across the software stack from model frameworks to kernels, contributing to graph optimizations and runtime improvements.

You will collaborate with compiler, runtime, and hardware teams, applying CUDA, ROCm, OpenCL, SYCL, or Triton expertise to improve transformer and LLM workloads while ensuring energy-efficient, scalable inference.

Qualifications

  • Bachelor’s degree in a technical field and 3–5 years of relevant experience, or a Master’s with 3+ years.
  • Strong foundations in algorithms, data structures, and systems programming.
  • Experience with Python and C/C++ in real-world projects or research.
  • Familiarity with deep learning concepts, transformer models, and GPU/accelerator architectures.

Responsibilities

  • Optimize deep learning computational graphs for performance, throughput, and energy efficiency on Ampere accelerators.
  • Develop graph-level optimizations such as fusion, memory planning, and quantization.
  • Analyze end-to-end performance across frameworks, runtimes, and hardware; build profiling infrastructure.
  • Collaborate with compiler, runtime, kernel, architecture, and hardware teams on co-design.

Skills

Python
C/C++
Graph algorithms
Systems programming
Profiling

Education

Bachelor's degree in Computer Science / Computer Engineering or related field

Tools

CUDA
ROCm
OpenCL
SYCL
Triton

Job description

Ampere is seeking a Software Engineer for its AI Accelerator team to optimize deep learning graphs and maximize performance on Ampere hardware. You will work across the software stack from model frameworks to kernels, contributing to graph optimizations and runtime improvements.

You will collaborate with compiler, runtime, and hardware teams, applying CUDA, ROCm, OpenCL, SYCL, or Triton expertise to improve transformer and LLM workloads while ensuring energy-efficient, scalable inference.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Compiler Engineer: Graph Optimizations & HW/SW Co-Design
AI Compiler Engineer: Graph Optimizations & HW/SW Co-Design

Ampere Computing • Santa Clara (CA)

Hybrid
USD 195,000 - 292,000
Health insurance
401K retirement plan
Unlimited flextime
+1
Senior AI Compiler Architect for Efficient Deep Learning
Senior AI Compiler Architect for Efficient Deep Learning

Ampere • Santa Clara (CA)

On-site
USD 195,000 - 292,000
Health insurance
401K plan
Unlimited flextime
AI Accelerator, Software Engineer- Graph Optimization/Compilers
AI Accelerator, Software Engineer- Graph Optimization/Compilers

Ampere • Santa Clara (CA)

On-site
USD 159,000 - 239,000
Health insurance
401K retirement plan
Unlimited flextime
Lead AI Graph Compiler Engineer
Lead AI Graph Compiler Engineer

EnCharge AI • United States

Remote
USD 190,000 - 255,000
Software Principal Engineer- AI Compiler
Software Principal Engineer- AI Compiler

Ampere Computing • Santa Clara (CA)

Hybrid
USD 195,000 - 292,000
Health insurance
401K retirement plan
Unlimited flextime
+1
Software Principal Engineer- AI Compiler
Software Principal Engineer- AI Compiler

Ampere • Santa Clara (CA)

On-site
USD 195,000 - 292,000
Health insurance
401K plan
Unlimited flextime
Senior AI/ML Systems Engineer for Accelerator Optimization
Senior AI/ML Systems Engineer for Accelerator Optimization

Socket.dev • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior AI Compiler Engineer - XLA & DL Graphs
Senior AI Compiler Engineer - XLA & DL Graphs

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
AI Accelerator Runtime Engineer — Execution & Performance
AI Accelerator Runtime Engineer — Execution & Performance

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Compiler Engineer (MLIR) — Equity & Impact
Senior AI Compiler Engineer (MLIR) — Equity & Impact

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000