Performance Modeling Architect

Morr0

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

15 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Morr0 is seeking a Principal Performance Modeling Architect in the Bay Area (onsite). You will own the performance analysis framework used to evaluate proposed silicon and system architectures before they are built.

You'll collaborate with senior GPU/AI architects to shape accelerator compute, memory, interconnect, and cluster-level systems, developing production-quality modeling tools in Python and C++. This hands-on IC role will influence RTL decisions and architecture choices, ensuring

Qualifications

  • Experience in GPU/TPU/NPU performance modeling.
  • Background in architecture simulation or analytical modeling for AI accelerators.
  • Experience with LLM training/inference performance.
  • Knowledge of memory hierarchy and bandwidth modeling.
  • Familiarity with NoC, fabric, and interconnect performance.
  • Experience with Python and C++ modeling frameworks.
  • Understanding roofline and bottleneck analysis.

Responsibilities

  • Build end-to-end performance models for AI accelerator platforms.
  • Model workloads across silicon, memory, interconnect, and multi-accelerator systems.
  • Translate proposals into projections of performance, power, and cost.
  • Identify bottlenecks and quantify architectural trade-offs before RTL/silicon.
  • Model LLM inference workloads including batching and KV-cache management.
  • Develop production-quality analytical frameworks in Python/C++.

Skills

Performance modeling
GPU/TPU/NPU modeling
Architecture analysis
LLM inference performance
System-level optimization

Tools

Python
C++

Job description

Principal Performance Modeling Architect | Bay Area | Onsite |

We’re working with an early-stage AI hardware startup building a next-gen compute platform spanning accelerator architecture, system software and large-scale AI infrastructure.

They’re hiring a Principal Performance Modeling Architect to own the performance analysis framework used to evaluate proposed silicon and system architectures before they are built.

This is a highly influential, hands-on IC role. Your work will directly shape architecture decisions across accelerator compute, memory, interconnect and cluster-level systems.

What you’ll own
  • Build and evolve end-to-end performance models for next-generation AI accelerator platforms
  • Model workloads across silicon, memory, interconnect and multi-accelerator systems
  • Translate architecture proposals into projections around performance, power/energy efficiency and cost
  • Identify system bottlenecks and quantify architectural trade-offs before RTL or silicon exists
  • Model modern LLM inference workloads including prefill/decode behaviour, KV cache, batching and parallelism
  • Work closely with senior GPU/AI architects to influence future silicon and system architecture
  • Develop production-quality analytical and simulation frameworks, primarily in Python/C++
What we’re looking for

Strong candidates are likely to come from GPU, TPU, NPU or AI accelerator teams at companies such as NVIDIA, Google, AMD, Intel, Meta, Microsoft, Amazon, Qualcomm, Cerebras, Groq or similar.

You’ll ideally have experience in several of the following:

  • GPU / TPU / NPU / AI accelerator performance modeling
  • Architecture simulation or analytical performance modeling
  • LLM training or inference performance
  • Memory hierarchy and bandwidth modeling
  • NoC, fabric or scale-up interconnect performance
  • Roofline, bottleneck or first-principles performance analysis
  • Python and/or C++ modeling frameworks

Experience with concepts such as vLLM, SGLang, KV-cache management, continuous batching, prefill/decode, tensor parallelism or pipeline parallelism would be particularly relevant.

Why it’s interesting

This isn’t a role where you inherit a mature simulator and optimise one component.

You’ll have broad ownership of the modeling platform and work alongside engineers defining the architecture itself, using performance analysis to answer some of the most important questions before the silicon is committed.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

On-site
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Seattle (WA)

On-site
USD 342,000 - 555,000
Performance Architect
Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Lead AI Accelerator Performance Architect
Lead AI Accelerator Performance Architect

Morr0 • San Francisco (CA)

On-site
USD 180,000 - 240,000
Performance Modeling Lead
Performance Modeling Lead

OpenAI • San Francisco (CA)

On-site
USD 342,000 - 555,000
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Los Angeles (CA)

On-site
USD 266,000 - 445,000
Performance Modeling Lead
Performance Modeling Lead

Showcify • United States

On-site
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Performance Modeling Engineer
Performance Modeling Engineer

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Relocation assistance
Performance Modeling Engineer
Performance Modeling Engineer

The Consensus • San Jose (CA)

On-site
USD 150,000 - 230,000
Medical, dental, and vision coverage
Housing subsidy near Santana Row: $2k/
Relocation support to San Jose
+3
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • United States

Hybrid
USD 90,000 - 130,000
Hybrid work model (3 days in office)
Relocation assistance