An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Morr0 is seeking a Principal Performance Modeling Architect in the Bay Area (onsite). You will own the performance analysis framework used to evaluate proposed silicon and system architectures before they are built.
You'll collaborate with senior GPU/AI architects to shape accelerator compute, memory, interconnect, and cluster-level systems, developing production-quality modeling tools in Python and C++. This hands-on IC role will influence RTL decisions and architecture choices, ensuring
We’re working with an early-stage AI hardware startup building a next-gen compute platform spanning accelerator architecture, system software and large-scale AI infrastructure.
They’re hiring a Principal Performance Modeling Architect to own the performance analysis framework used to evaluate proposed silicon and system architectures before they are built.
This is a highly influential, hands-on IC role. Your work will directly shape architecture decisions across accelerator compute, memory, interconnect and cluster-level systems.
Strong candidates are likely to come from GPU, TPU, NPU or AI accelerator teams at companies such as NVIDIA, Google, AMD, Intel, Meta, Microsoft, Amazon, Qualcomm, Cerebras, Groq or similar.
You’ll ideally have experience in several of the following:
Experience with concepts such as vLLM, SGLang, KV-cache management, continuous batching, prefill/decode, tensor parallelism or pipeline parallelism would be particularly relevant.
This isn’t a role where you inherit a mature simulator and optimise one component.
You’ll have broad ownership of the modeling platform and work alongside engineers defining the architecture itself, using performance analysis to answer some of the most important questions before the silicon is committed.