Senior AI Inference Runtime Architect

Arm

Seattle (WA)

Hybrid

USD 263,000 - 355,000

Full time

21 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid working
Accommodations during recruitment

Job summary

Arm is seeking a Principal Software Engineer to lead the AI Inference Runtime team in Seattle, shaping technical direction for distributed inference workloads and memory systems. You will collaborate with AI Infra, compute, and product groups to optimize performance across platforms and models.

You will architect scheduling, batching, KV‑cache management, and kernel development, driving end‑to‑end model support and production validation with a focus on latency, throughput, and resource

Qualifications

  • 8+ years in ML or high‑performance systems, or equivalent impact.
  • Deep AI inference knowledge including model execution and KV-cache behavior.
  • Strong coding skills in C++, Rust, or Python with concurrency and memory handling.

Responsibilities

  • Define architecture and roadmap for AI inference runtime across platforms.
  • Enable end‑to‑end model support with scheduling, batching, memory management.
  • Profile bottlenecks and develop optimized kernels and data paths.
  • Lead reviews, mentor engineers, and drive performance‑engineering practices.

Skills

ML systems
C++/Rust/Python
Profiling/optimization
Distributed systems

Job description

Arm is seeking a Principal Software Engineer to lead the AI Inference Runtime team in Seattle, shaping technical direction for distributed inference workloads and memory systems. You will collaborate with AI Infra, compute, and product groups to optimize performance across platforms and models.

You will architect scheduling, batching, KV‑cache management, and kernel development, driving end‑to‑end model support and production validation with a focus on latency, throughput, and resource

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Inference Runtime Architect
Principal AI Inference Runtime Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Staff AI Inference Runtime Engineer
Staff AI Inference Runtime Engineer

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Hybrid working
Recruitment accommodations
Senior AI Inference Runtime Engineer - Distributed
Senior AI Inference Runtime Engineer - Distributed

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Senior AI Compute Infra Engineer
Senior AI Compute Infra Engineer

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Senior AI Compute Platform Engineer
Senior AI Compute Platform Engineer

Arm • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Senior AI Compute Infrastructure Architect
Senior AI Compute Infrastructure Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Relocation package
Visa sponsorship
Principal AI Inference Cloud Architect
Principal AI Inference Cloud Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Staff Engineer, Inference Runtime — High-Performance AI Serving
Staff Engineer, Inference Runtime — High-Performance AI Serving

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Senior Cloud AI Inference Engineer
Senior Cloud AI Inference Engineer

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Visa sponsorship
Cloud AI Inference Platform Architect
Cloud AI Inference Platform Architect

Arm • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Relocation package with visa Spons.&
Hybrid working options