AI Accelerator Software Principal Engineer - Inference

Ampere

Portland (OR)

On-site

USD 182,000 - 273,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) retirement plan
Unlimited flextime
10+ paid holidays

Job summary

Ampere is seeking an AI Accelerator Software Principal Engineer – NPU Full‑Stack Integration to lead high‑performance, low‑latency inference on Arm Ethos U85. You will optimize software stacks from model execution to accelerator kernels and collaborate with hardware teams for scalable AI workloads.

The role demands strong C/C++, Python, and Linux systems expertise, plus ML understanding and modern AI tooling fluency. Expect a competitive base pay and comprehensive benefits.

Qualifications

  • BS in CS/CE/EE or related field with 8+ years experience; or MS 6+ years; or PhD 3+ years.
  • Experience with AOT compilation in PyTorch and edge deployment paths.
  • Linux accelerator/runtime development experience preferred.

Responsibilities

  • End-to-end deep learning performance acceleration.
  • Model enablement with quality and speed for edge devices.
  • Hardware/software co-design and optimization with platform teams.
  • Build AI software components and optimized execution paths.
  • Cross-functional collaboration with compiler, kernels, and product teams.

Skills

Python
C/C++
Performance profiling
Linux development
AI/ML concepts

Education

Bachelor's degree in CS/CE/EE
Master's degree
PhD

Tools

PyTorch
Compiler/runtime
Linux kernel

Job description

Ampere is seeking an AI Accelerator Software Principal Engineer – NPU Full‑Stack Integration to lead high‑performance, low‑latency inference on Arm Ethos U85. You will optimize software stacks from model execution to accelerator kernels and collaborate with hardware teams for scalable AI workloads.

The role demands strong C/C++, Python, and Linux systems expertise, plus ML understanding and modern AI tooling fluency. Expect a competitive base pay and comprehensive benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Accelerator Software Principal Engineer Fast Inference
AI Accelerator Software Principal Engineer Fast Inference

Ampere • Portland (OR)

On-site
USD 182,000 - 273,000
Premium medical, dental, and vision insurance
Unlimited Flextime
401K retirement plan
AI Accelerator Software Principal Engineer – NPU Full-Stack Integration
AI Accelerator Software Principal Engineer – NPU Full-Stack Integration

Ampere • Portland (OR)

On-site
USD 182,000 - 273,000
Health insurance
401(k) retirement plan
Unlimited flextime
+1
AI Accelerator Software Principal Engineer- Framework Integration
AI Accelerator Software Principal Engineer- Framework Integration

Ampere • Portland (OR)

On-site
USD 182,000 - 273,000
Premium medical, dental, and vision insurance
Unlimited Flextime
401K retirement plan
Senior AI Compiler Architect for Efficient Deep Learning
Senior AI Compiler Architect for Efficient Deep Learning

Ampere • Santa Clara (CA)

On-site
USD 195,000 - 292,000
Health insurance
401K plan
Unlimited flextime
AI Systems Engineer: High-Performance ML on Accelerators
AI Systems Engineer: High-Performance ML on Accelerators

AMD • San Jose (CA)

On-site
USD 150,000 - 190,000
Software Principal Engineer- AI Compiler
Software Principal Engineer- AI Compiler

Ampere Computing • Santa Clara (CA)

Hybrid
USD 195,000 - 292,000
Health insurance
401K retirement plan
Unlimited flextime
+1
AI Systems Engineer: ML Kernels & HPC Acceleration
AI Systems Engineer: ML Kernels & HPC Acceleration

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Principal MLIR Compiler Architect for AI Acceleration
Principal MLIR Compiler Architect for AI Acceleration

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 240,000
Software Principal Engineer- AI Compiler
Software Principal Engineer- AI Compiler

Ampere • Santa Clara (CA)

On-site
USD 195,000 - 292,000
Health insurance
401K plan
Unlimited flextime
AI Performance Engineer – HPC, ARM & Distributed Inference
AI Performance Engineer – HPC, ARM & Distributed Inference

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000