Lead AI Systems Architect - MoE Runtime & Memory Hierarchy

Lexarenterprise

San Jose (CA)

Hybrid

USD 210,000 - 350,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Lexar Enterprise seeks a senior systems architect to own the end-to-end technical architecture for sparse MoE inference and memory management. You will convert strategic direction into concrete models, APIs, data paths, cache policies, and performance models across runtimes and hardware.

You will bridge model architecture, runtime software, OS memory management, and accelerator execution, guiding design from prototype to production while ensuring latency, reliability, and security across mobile,

Qualifications

  • 15+ years in systems software, storage software, HPC, OS, ML infra, or related engineering.
  • BS/MS degree in Computer Science or equivalent experience
  • At least 5 years in AI/ML infrastructure, transformer inference, model serving, or accelerator software.
  • Demonstrated experience architecting complex systems across multiple software or hardware layers.
  • Strong C++ and Python skills with debugging concurrency, memory, I/O, and performance issues.
  • Deep Linux systems knowledge (VMs, mmap, page cache, NUMA, DMA, AIO).
  • Experience with NVMe/UFS/PCIe or other storage interfaces and firmware.
  • Ability to translate architecture into testable requirements and milestones.

Responsibilities

  • Architect the model-aware residency system with explicit lifecycle, identity, placement, quality, deadline, and security attributes.
  • Design HLC policies for hot, warm, cold, and prefetched states including eviction, routing telemetry, and compression decisions.
  • Develop quantitative workload models for tokens, cache hits, throughput, and bandwidth.
  • Own deployment tracks for mobile/edge and enterprise/near-data with diverse storage and execution paths.
  • Integrate with runtimes like llama.cpp, vLLM, SGLang, ExecuTorch, and vendor stacks; determine OS vs. hardware changes.
  • Architect agent-state storage across accelerator memory, DRAM, CXL, local SSD, and remote storage.
  • Design runtime-to-system APIs for expert requests, prefetch deadlines, and telemetry.
  • Lead performance and correctness reviews to balance latency, energy, and cost without compromising correctness.
  • Guide implementation, reviewing C++, Python, kernel, firmware, and simulator designs; mentor senior engineers.

Skills

C++
Python
Linux systems
Transformer inference
MoE routing
Performance tuning
Memory management
GPU/CPU/NPU
Software architecture
Kernel/firmware

Education

BS/MS in Computer Science

Tools

llama.cpp
vLLM
SGLang
ExecuTorch

Job description

Lexar Enterprise seeks a senior systems architect to own the end-to-end technical architecture for sparse MoE inference and memory management. You will convert strategic direction into concrete models, APIs, data paths, cache policies, and performance models across runtimes and hardware.

You will bridge model architecture, runtime software, OS memory management, and accelerator execution, guiding design from prototype to production while ensuring latency, reliability, and security across mobile,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Systems Architect, MoE Runtime & Memory Hierarchy
Principal AI Systems Architect, MoE Runtime & Memory Hierarchy

Lexarenterprise • San Jose (CA)

Hybrid
USD 210,000 - 350,000
Senior Architect – Android Tiered Memory Systems
Senior Architect – Android Tiered Memory Systems

Lexarenterprise • San Jose (CA)

Hybrid
USD 170,000 - 240,000
Lead AI Architect - Memory & Multi-Agent Systems
Lead AI Architect - Memory & Multi-Agent Systems

EY • Arlington (VA)

Hybrid
USD 144,000 - 329,000
Lead AI Architect - Memory & Multi-Agent Systems
Lead AI Architect - Memory & Multi-Agent Systems

EY • Memphis (TN)

Hybrid
USD 144,000 - 329,000
Medical and dental coverage
401(k) & pension
Paid time off
+2
Lead AI Architect: Memory-Driven Multi-Agent Systems
Lead AI Architect: Memory-Driven Multi-Agent Systems

EY • Huntsville (AL)

On-site
USD 144,000 - 329,000
Medical & dental coverage
401(k) plans
Paid time off
+2
Technical Director, Large-Scale AI Model Inferencing
Technical Director, Large-Scale AI Model Inferencing

samsungsemiconductor • San Jose (CA)

On-site
USD 180,000 - 240,000
Lead AI Memory Architect & Systems Design Manager
Lead AI Memory Architect & Systems Design Manager

EY • Westlake Village (CA)

Hybrid
USD 173,000 - 374,000
Hybrid work model
Medical & dental coverage + 401(k)
Flexible vacation policy
Lead AI Architect: Memory & Multi-Agent Systems
Lead AI Architect: Memory & Multi-Agent Systems

EY • Dallas (TX)

Hybrid
USD 144,000 - 329,000
Lead AI Architect: Memory Systems & Multi-Agent AI
Lead AI Architect: Memory Systems & Multi-Agent AI

EY • St. Louis (MO)

Hybrid
USD 144,000 - 329,000
Hybrid work model
Total Rewards package
Medical and dental coverage
+2
Principal AI Memory Architect
Principal AI Memory Architect

Conductor • San Jose (CA)

On-site
USD 219,000 - 351,000
Medical/Dental/Vision/401k
4+ weeks paid time off
Fertility/adoption support