Senior Lead AI Systems and Performance Architect – LLM Inference

Intel Corporation

Prescott (AZ)

Hybrid

USD 134,000 - 179,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Intel Corporation is seeking a Senior Lead AI Systems and Performance Architect to drive the End-to-End LLM Inference Simulation Platform. You will lead the simulation platform, define modeling workflows, and conduct dynamic trace replays to identify full-system bottlenecks and optimize GPU usage, memory sizing, and interconnects.

The role involves coordinating with silicon architects and software teams, mentoring engineers, and delivering actionable insights for Intel's AI inference SoC and

Qualifications

  • Ph.D. or Master’s in CS/CE/EE or related quantitative field.
  • 4–6+ years (PhD) or 6–8+ years (Master’s) or 8–10+ years (Bachelor) in system performance modeling or AI systems.
  • Hands-on experience with discrete-event, trace-driven, or analytical performance simulators.
  • Proven ability to identify bottlenecks across compute, memory, and interconnects in multi-accelerator environments.
  • Proficiency in Python and C++ (14/17/20)

Responsibilities

  • Lead the technical direction of a discrete-event simulation platform for multi-accelerator inference systems.
  • Model serving mechanisms and interconnects, and automate simulation pipelines.
  • Perform bottleneck analysis across compute, memory, and interconnects; mentor team.

Skills

Python
C++
Discrete-event simulation
Performance modeling
System performance

Education

Ph.D. or Master's in CS/CE/EE or related field

Tools

Simulation frameworks
Profiling tools
Automation scripting

Job description

Job Details:
Job Description:

About the Role and TeamThe Intel DCG AISOC organization is developing the future of high-performance accelerated AI platforms. Real-world inference performance, latency SLOs, and TCO are governed by full-system interplay: dynamic batching, distributed KV cache movement, scale-up fabrics, and heterogeneous host-accelerator memory hierarchies. We are seeking a hands-on, technically deep Senior Lead AI Systems and Performance Architect to drive the technical direction, development, and system-level analysis of our End-to-End LLM Inference Simulation Platform. In this role, you will lead the technical execution of the simulation platform, define modeling workflows, and conduct dynamic trace replays. Your mission is to identify full-system bottlenecks, quantify GPU optimization, evaluate host CPU offloading, size memory subsystems, and assess high-speed interconnects – translating simulation findings into actionable engineering insights to guide Intel's future AI inference SoC and platform architectures.

Key Responsibilities
1. Simulation Platform and Workflow Leadership
  • Technical Ownership: Lead the technical direction and continuous evolution of a discrete-event simulation platform modeling distributed multi-accelerator inference systems.
  • Serving Stack Modeling: Model critical serving mechanisms: e.g. continuous batching, chunked prefill, prefill-decode disaggregation, and dynamic KV cache allocation.
  • Fabric and Pipeline Abstraction: Integrate behavioral models for scale-up interconnects (UALink, PCIe Gen 6/7), capturing latency, arbitration, and collective communications (All-to-All, All-Reduce).
  • Workflow Automation: Design automated simulation pipelines to ingest empirical accelerator performance tables, execute parameter sweeps, and analyze results.
2. Dynamic Trace Replay and Bottleneck Analysis
  • Production Trace Replay: Drive the ingestion and replay of realistic production traces, capturing arrival burstiness, diverse prompt/output distributions, and shared prefix contexts.
  • Agentic Workload Modeling: Keep the simulator aligned with emerging patterns, specifically Agentic workflows (multi-step tool calls, state branches) and long-context KV cache demands.
  • Root-Cause Isolation: Conduct sensitivity analyses to isolate bottlenecks across compute throughput, memory bandwidth, bus latency, scale-up fabric contention, or scheduling overheads.
  • Metric Evaluation: Evaluate trade-offs across SLOs (TTFT, TPOT, ITL, queue wait times) against cluster throughput, power, and cost.
3. HW/SW Co-Design and Platform Recommendations
  • Silicon Architecture Input: Translate findings into architectural proposals for the AI SoC team, defining balanced compute-to-memory ratios, SRAM sizes, and interface configurations.
  • Platform Optimization: Evaluate the host CPU's role in the inference pipeline to identify offloading opportunities (orchestration, tokenization, routing, tiered KV cache staging).
  • Memory Hierarchy Sizing: Provide workload-grounded capacity and bandwidth recommendations, evaluating HBM/LPDDR vs. tiered system memory.
4. Team Leadership and Collaboration
  • Mentorship and Quality: Guide team members on simulation methodologies, code quality, model calibration, and experimental rigor.
  • Cross-Team Alignment: Partner with silicon architects, software engineers, and platform teams to translate architectural questions into structured simulation studies.
Qualifications
Minimum Qualifications
  • Education: Ph.D. or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related quantitative field.
  • Experience: Ph.D. with 4-6+ years (or Master's with 6-8+ years, or Bachelor's with 8-10+ years) in system performance modeling, computer architecture, or AI systems engineering.
  • Simulation and Modeling: Hands-on experience developing or utilizing discrete-event, trace-driven, or analytical performance simulators.
  • Performance Analysis: Proven track record of identifying performance bottlenecks across compute, memory, and interconnects in multi-accelerator/heterogeneous environments.
  • Programming: Proficiency in Python and C++ (14/17/20), with experience in modular software design and automated data processing.
Preferred Qualifications
  • AI Serving Frameworks: Understanding of LLM serving runtimes (e.g., vLLM, TensorRT-LLM, SGLang) and mechanisms (PagedAttention, continuous batching, speculative decoding).
  • Interconnect and Fabrics: Knowledge of accelerator interconnects, particularly UALink, PCIe Gen 6/7, CXL, or high-performance networking (RoCEv2, Ultra Ethernet).
  • Heterogeneous Co-Design: Experience evaluating host-accelerator partitioning, CPU-assisted pipelines, or tiered memory architectures.
  • Workload Characterization: Experience profiling and replaying production trace telemetry (multi-turn conversations, MoE routing, Agentic workflows).
  • Technical Leadership: Track record of leading technical projects, driving team consensus, and mentoring engineers.
Job Type

College Grad

Shift

Shift 1 (China)

Primary Location

PRC, Shanghai

Posting Statement

All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation or any other characteristic protected by local law, regulation, or ordinance.

Position of Trust

N/A

Work Model for this Role

This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. Job posting details (such as work model, location or time type) are subject to change.

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Systems Performance and Simulation Engineer
AI Systems Performance and Simulation Engineer

Intel Corporation • Prescott (AZ)

On-site
USD 65,000 - 90,000
Software Engineer — Distributed LLM Inference Systems
Software Engineer — Distributed LLM Inference Systems

Intel Corporation • Prescott (AZ)

On-site
USD 52,000 - 82,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
AI Software Technical Intern
AI Software Technical Intern

Intel • Santa Clara (CA)

Hybrid
USD 133,000 - 171,000
Hybrid work model
Competitive internship compensation
Mentorship from senior engineers
•AI/ML CAD Engineer, RTL & Design Verification
•AI/ML CAD Engineer, RTL & Design Verification

Intel • Austin (TX)

Hybrid
USD 122,000 - 232,000
AI Framework Software Engineer - vLLM
AI Framework Software Engineer - vLLM

Intel Corporation • Prescott (AZ)

On-site
USD 45,000 - 81,000
Platform Power Thermal Performance Engineer
Platform Power Thermal Performance Engineer

Intel Corporation • Prescott (AZ)

On-site
USD 27,000 - 36,000
AI Frameworks Engineer
AI Frameworks Engineer

Intel Corporation • Santa Clara (CA)

On-site
USD 150,000 - 276,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Research Engineer/Scientist
AI Research Engineer/Scientist

Intel Corporation • Prescott (AZ)

On-site
USD 52,000 - 82,000
AI Software Solutions Engineer
AI Software Solutions Engineer

Intel Corporation • Prescott (AZ)

On-site
USD 39,000 - 63,000