Senior AI Inference Runtime Engineer - Distributed

Arm Limited

Seattle (WA)

Hybrid

USD 209,000 - 283,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Arm Limited is seeking a Software Engineer for the AI Inference Runtime team, driving architecture and roadmaps for distributed inference workloads and model execution across hardware platforms.

You will lead profiling, kernel optimization, and efficiency improvements, collaborating with cloud, framework, compiler, hardware, and research teams. The role involves hands-on implementation, performance engineering, and mentoring engineers within Arm's AI Platforms group.

Qualifications

  • 5+ years of experience in ML systems, high-performance systems, compilers, or production AI inference.
  • Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.
  • Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, memory management, and data movement.
  • Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, OS, and hardware.

Responsibilities

  • Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms.
  • Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.
  • Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.
  • Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.
  • Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices.

Skills

5+ years experience
In-depth AI inference
C++, Rust, Python
Profiling & optimization

Job description

Arm Limited is seeking a Software Engineer for the AI Inference Runtime team, driving architecture and roadmaps for distributed inference workloads and model execution across hardware platforms.

You will lead profiling, kernel optimization, and efficiency improvements, collaborating with cloud, framework, compiler, hardware, and research teams. The role involves hands-on implementation, performance engineering, and mentoring engineers within Arm's AI Platforms group.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Inference Runtime Architect
Principal AI Inference Runtime Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Senior Cloud AI Inference Engineer
Senior Cloud AI Inference Engineer

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Visa sponsorship
Principal AI Inference Cloud Architect
Principal AI Inference Cloud Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Senior AI Compute Infrastructure Architect
Senior AI Compute Infrastructure Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Relocation package
Visa sponsorship
Principal Software Engineer, AI Inference Runtime
Principal Software Engineer, AI Inference Runtime

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
AI Performance Engineer – HPC, ARM & Distributed Inference
AI Performance Engineer – HPC, ARM & Distributed Inference

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000
Staff Software Engineer, AI Inference Runtime
Staff Software Engineer, AI Inference Runtime

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Platform Engineer for AI Inference & Optimization
Platform Engineer for AI Inference & Optimization

OpenAI • Seattle (WA)

On-site
USD 180,000 - 240,000
Distributed Inference Performance Engineer
Distributed Inference Performance Engineer

OpenAI • California (MO)

On-site
USD 150,000 - 190,000
Model Runtime Engineer for Frontier AI Inference
Model Runtime Engineer for Frontier AI Inference

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000