Staff AI Inference Runtime Engineer

Arm

Seattle (WA)

Hybrid

USD 209,000 - 283,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid working
Recruitment accommodations

Job summary

Arm in Seattle is seeking a Software Engineer for the AI Inference Runtime team to set technical direction for distributed inference runtime powering SOTA AI models. You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed execution, kernel development and optimization, shaping how efficiently models use compute.

You will partner with AI Infrastructure, compute and product teams to raise performance and energy efficiency of Arm’s AI platform,

Qualifications

  • 5+ years in ML systems, high-performance systems, compilers, kernel development, or production AI inference.
  • Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.
  • Strong programming in C++, Rust, Python with concurrency and memory management.
  • Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, and hardware.

Responsibilities

  • Define architecture, interfaces, and roadmap for AI inference runtime capabilities.
  • Enable model architectures end to end through operator support, validation, and optimization of scheduling, batching, and memory management.
  • Profile bottlenecks and develop optimized kernels and data movement paths across compute, memory, networking, and framework integration.
  • Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems.
  • Partner with cloud, framework, compiler, hardware, and research teams; mentor engineers and drive performance-engineering practices.

Skills

ML systems
high-performance systems
compilers
kernel development
performance optimization
concurrency & parallelism
memory management
profiling & debugging

Tools

C++
Rust
Python
Assembler/Intrinsics

Job description

Arm in Seattle is seeking a Software Engineer for the AI Inference Runtime team to set technical direction for distributed inference runtime powering SOTA AI models. You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed execution, kernel development and optimization, shaping how efficiently models use compute.

You will partner with AI Infrastructure, compute and product teams to raise performance and energy efficiency of Arm’s AI platform,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Inference Runtime Architect
Principal AI Inference Runtime Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Senior AI Inference Runtime Engineer - Distributed
Senior AI Inference Runtime Engineer - Distributed

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Senior AI Compute Infra Engineer
Senior AI Compute Infra Engineer

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Senior Cloud AI Inference Engineer
Senior Cloud AI Inference Engineer

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Visa sponsorship
Staff Software Engineer, AI Inference Runtime
Staff Software Engineer, AI Inference Runtime

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Hybrid working
Recruitment accommodations
Senior AI Compute Platform Engineer
Senior AI Compute Platform Engineer

Arm • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Staff Software Engineer, AI Inference Runtime
Staff Software Engineer, AI Inference Runtime

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Principal Software Engineer, AI Inference Runtime
Principal Software Engineer, AI Inference Runtime

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Senior AI Compute Infrastructure Architect
Senior AI Compute Infrastructure Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Relocation package
Visa sponsorship
Staff Engineer, Inference Runtime — High-Performance AI Serving
Staff Engineer, Inference Runtime — High-Performance AI Serving

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000