Senior Performance Modeling Architect, CPU Fabric and LLC

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 288,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

NVIDIA Corporation in Santa Clara is seeking a Performance Modeling Architect to define and improve CPU cache hierarchies and interconnects for automotive and data center systems.

You will develop cycle-accurate performance models (C++/SystemC), analyze bottlenecks, and evaluate coherency protocols. Collaboration with software and verification teams will optimize drivers and validate models against silicon.

Qualifications

  • Master’s or Ph.D. in Computer Engineering, Electrical Engineering, or Computer Science (or equivalent) with 5+ years of experience.
  • Strong understanding of CPU microarchitecture, memory consistency models, and cache coherency protocols.
  • Proven experience in C++ or SystemC for cycle-accurate or functional modeling.
  • Proficiency in Python or similar scripting languages for data processing and visualization.
  • Understanding of NoC topologies (Mesh, Ring, Torus) and arbitration logic.

Responsibilities

  • Develop and maintain high-fidelity, cycle-accurate performance models for coherent interconnects and large-scale caches.
  • Model and analyze performance bottlenecks across automotive and data center scales.
  • Evaluate performance impact of coherency protocols and snooping filters.
  • Run industry benchmarks to drive architectural trade-offs and optimize drivers.

Skills

CPU microarchitecture
Memory coherency
C++/SystemC
Python scripting
NoC topologies

Education

Master's or PhD in relevant field

Tools

C++
SystemC

Job description

We are looking for a highly skilled Performance Modeling Architect to lead the architectural definition and improvement of our next-generation CPU Cache Hierarchies and interconnects. This is an outstanding chance to create scalable solutions that connect two fast-paced domains: the high-reliability, low-latency needs of Automotive and the massive efficiency, high-density demands of Data Center systems. You will build the "source of truth" models that govern data movement across our silicon, ensuring our next-level caches (L3/System Cache) and coherent fabrics achieve ambitious performance goals.

What you'll be doing

As a core member of the architecture team, your daily work will involve: Developing and maintaining high-fidelity, cycle-accurate performance models (C++/SystemC) for coherent interconnects and large-scale shared caches. Modeling and analyzing performance bottlenecks across varying scales, from small-cluster automotive SoCs to massive, multi-mesh data center architectures. Evaluating the performance impact of different coherency protocols (e.g., CHI, ACE, or proprietary) and snooping filters. Running and analyzing industry-standard benchmarks (SPEC, MLPerf, Automotive-specific suites) to drive architectural trade-offs. Collaborating with build and verification teams to correlate performance models with silicon and working with software teams to optimize drivers for the underlying hardware topology.

What we need to see
  • A Master’s or Ph.D. in Computer Engineering, Electrical Engineering, or Computer Science (or equivalent experience) with a focus on architecture with 5+ years of experience.
  • Strong understanding of CPU microarchitecture, memory consistency models, and cache coherency protocols.
  • Proven experience in C++ or SystemC for cycle-accurate or functional modeling.
  • Proficiency in Python or similar scripting languages for processing large datasets, generating performance visualizations, and automating simulation sweeps.
  • Understanding of Network-on-Chip (NoC) topologies (Mesh, Ring, Torus), credit-based flow control, and arbitration logic.
Ways to stand out from the crowd
  • Cross-Domain Versatility: Practical experience managing the functional safety (ISO 26262) requirements of automotive chips alongside the power-performance-area (PPA) limitations of data center hardware.
  • Hardware Performance Counters: Experience defining or using PMU (Performance Monitoring Unit) events to debug performance on real silicon or emulators.
  • Formal Methods: A background in using formal verification or mathematical modeling to prove the correctness of complex coherency state machines.
  • Custom Tooling: A history of building your own internal tools or frameworks to accelerate architectural exploration rather than just using off-the-shelf simulators.
  • Advanced Memory Systems: Knowledge of emerging memory technologies like CXL (Compute Express Link) or HBM (High Bandwidth Memory) and how they collaborate with coherent fabrics.
Your base salary

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Architect, SoC and Systems Modelling
Principal Architect, SoC and Systems Modelling

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Senior Cache Coherency Architect
Senior Cache Coherency Architect

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity and benefits
Senior System Simulation Architect
Senior System Simulation Architect

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Distinguished Engineer, End-to-End Scaling Performance Architecture
Distinguished Engineer, End-to-End Scaling Performance Architecture

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Senior Performance Verification Engineer
Senior Performance Verification Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 136,000 - 265,000
Equity
Benefits
Senior System Simulation Architect
Senior System Simulation Architect

NVIDIA • Durham (NC)

On-site
USD 184,000 - 287,500
Equity
Benefits
CPU Performance Architect
CPU Performance Architect

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Architect, GPU and SoC Modeling
Senior Architect, GPU and SoC Modeling

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Logic Design Engineer, Cache Coherent Interconnects
Senior Logic Design Engineer, Cache Coherent Interconnects

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 136,000 - 265,000
Principal CPU Performance Architect
Principal CPU Performance Architect

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits package