Senior AI Performance Engineer — Edge Inference Expert

Arm

San Jose (CA)

Hybrid

USD 263,000 - 355,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Arm is seeking an engineer in San Jose to optimize AI workloads on Arm technology, delivering best-in-class inference performance for production models on edge devices. The role involves kernel-level to system-level optimization, frequent collaboration with customers across the Bay Area, and influencing tooling roadmaps with real-world insights.

Ideal candidates have deep knowledge of DNN optimization, parallel memory hierarchies, and strong Python/C++ development skills, with experience in

Qualifications

  • Experience optimizing DNNs in Triton, CUDA or other kernel level programming language.
  • Deep understanding of parallel computing, memory hierarchies and performance optimization techniques for DNNs.
  • Strong programming skills in Python and C++, with solid experience in modern AI frameworks and execution models, plus profiling/analysis tools.
  • Strong communication and interpersonal skills.

Responsibilities

  • Develop highly optimized solutions for AI workloads, from kernel level to system level, to meet the needs of the customer application.
  • Create production quality reference implementations, documentation, and performance focused technical content
  • Act as a technical bridge between customers and internal teams, driving resolution of complex performance issues
  • Influence Arm’s IP and software roadmap through insights from real-world customer use cases

Skills

DNN optimization
CUDA
Kernel programming
Python
C++

Tools

Triton
CUDA

Job description

Arm is seeking an engineer in San Jose to optimize AI workloads on Arm technology, delivering best-in-class inference performance for production models on edge devices. The role involves kernel-level to system-level optimization, frequent collaboration with customers across the Bay Area, and influencing tooling roadmaps with real-world insights.

Ideal candidates have deep knowledge of DNN optimization, parallel memory hierarchies, and strong Python/C++ development skills, with experience in

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Performance Engineer - Edge Inference
Senior AI Performance Engineer - Edge Inference

Arm Limited • San Jose (CA)

On-site
USD 263,000 - 355,000
AI Inference Engineer - Kernel & Edge Optimization
AI Inference Engineer - Kernel & Edge Optimization

Jobgether SRL • United States

Remote
USD 82,000 - 177,000
Remote-first team
Exposure to cutting-edge AI research
Collaborative engineering environment
AI Performance Engineer – HPC, ARM & Distributed Inference
AI Performance Engineer – HPC, ARM & Distributed Inference

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000
Senior AI Inference Runtime Architect
Senior AI Inference Runtime Architect

Arm • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Senior AI Inference Runtime Engineer - Distributed
Senior AI Inference Runtime Engineer - Distributed

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Staff AI Inference Runtime Engineer
Staff AI Inference Runtime Engineer

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Hybrid working
Recruitment accommodations
Edge AI Kernel Engineer - High-Performance Inference
Edge AI Kernel Engineer - High-Performance Inference

QUADRIC PTY LTD • Burlingame (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
Company-paid life insurance
Voluntary life insurance
+11
Senior AI Kernel Engineer - Edge ML Optimization
Senior AI Kernel Engineer - Edge ML Optimization

Quadric Inc. • Burlingame (CA)

Hybrid
USD 110,000 - 270,000
Competitive salary
Equity
Health, dental, and vision
+7
AI Inference Engineer: Kernel & Edge Optimization
AI Inference Engineer: Kernel & Edge Optimization

Lever, Inc. • Town of Italy (NY)

On-site
EUR 90,000 - 130,000
Remote-first environment
International team
Exposure to cutting-edge AI research
+1
Senior SoC Performance Architect — AI/HPC Benchmarking
Senior SoC Performance Architect — AI/HPC Benchmarking

Arm • San Jose (CA)

On-site
USD 262,700 - 355,400