Performance Modeling Architect - AI Memory Systems

CyberCoders

Santa Clara (CA)

On-site

USD 200,000 - 250,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Relocation assistance and visa sponsor
Daily lunch stipend
Equity grant
Health, dental, vision, and life
401k

Job summary

CyberCoders in Santa Clara, CA (with Boston, MA option) seeks a Member of Technical Staff, Performance Modeling to develop models for a fabric-attached memory expansion device for AI accelerators. You will collaborate with silicon architects and workload teams to explore tradeoffs, validate assumptions, and uncover bottlenecks early in design.

The role emphasizes reasoning from first principles, working with incomplete information, and co-exploring the design space as hardware and software

Qualifications

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or a closely related field.
  • 5–10+ years of experience in performance modeling for data movement devices (NICs, memory expansion cards like CXL, IPU/DPU, NoC).
  • Ability to learn new ML architectures quickly and build performance models for them.
  • Ability to reason across multiple abstraction layers, from architecture to system-level performance.

Responsibilities

  • Build and maintain system-level performance models for a high-bandwidth data movement device in the scale-up domain.
  • Model workload from software memory access patterns to data distribution in the network and to on-device memory channels.
  • Collaborate with silicon architects, system designers, and workload owners to align performance expectations.
  • Identify bottlenecks, scaling limits, and sensitivity points across compute, memory, and interconnects in end-to-end workloads.
  • Clearly communicate modeling assumptions, limitations, and conclusions to technical and non-specialist stakeholders.

Skills

AI Memory Systems
Performance Modeling
Memory Expansion
Memory Systems Architecture
ML Systems
CUDA Memory Management

Education

Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or a closely related field

Tools

CUDA

Job description

Job Title: Performance Modeling Architect - AI Memory Systems

Job Location: Santa Clara, CA, or Boston, MA

Compensation: $200K - $250K base DOE plus 25% bonus and meaningful equity

Requirements: AI Memory Systems, Performance Modeling, Memory Expansion (NICs, SmartNICs, CXL, IPU/DPU, NoC), Memory Systems Architecture, ML Systems, CUDA Memory Management

Position Overview

We are seeking a Member of Technical Staff, Performance Modeling to develop performance models for our fabric-attached memory expansion device for AI accelerators. You'll work closely with silicon architects and workload teams to explore design tradeoffs, validate performance assumptions, and identify bottlenecks early in the development cycle. This role is well-suited for engineers who enjoy reasoning from first principles, working with incomplete information, and co-exploring the design space as hardware and software evolve together.

Key Responsibilities
  • Build and maintain system-level performance models for a high-bandwidth data movement device operating in the scale-up domain.
  • Model workload from software memory access patterns to data distribution in the network and all the way down to on-device memory channels.
  • Work day-to-day with silicon architects, system designers, and workload owners to align performance expectations and constraints.
  • Identify performance bottlenecks, scaling limits, and sensitivity points across compute, memory, and interconnects in end-to-end workload settings.
  • Clearly communicate modeling assumptions, limitations, and conclusions to both technical and non-specialist stakeholders.
Qualifications
  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or a closely related field.
  • Ability to quickly learn new ML architectures as soon as they come out, and build performance models for them.
  • 5-10+ years of experience in performance modeling for data movement devices: NICs, memory expansion cards (e.g., CXL), IPU/DPU, NoC.
  • Ability to reason across multiple abstraction layers, from architectural details to system-level performance behavior.
Preferred Qualifications
  • PhD in Computer Science, Electrical Engineering, or a related field.
  • Prior experience modeling performance for networking protocols with memory semantics.
  • Understanding of ML systems: workload sharding, KV caching hierarchies, attention optimizations, trade-offs when deploying ML models at scale, and various assumptions.
  • Familiarity with shared memory systems and frameworks (e.g., CUDA VMM).
  • Experience with scale-up and high-bandwidth interconnects (e.g., NVLink or similar technologies).
Benefits
  • Competitive salary commensurate with experience including base salary, performance-based bonus, and early-stage equity grant
  • Comprehensive benefits including health, dental, vision, and life insurance
  • Well-equipped, sunny offices in Santa Clara, CA and Boston, MA
  • Relocation assistance and visa sponsorship
  • Perks include a daily lunch stipend, 401k match, and more
  • A collaborative, continuous-learning work environment with smart, dedicated colleagues engaged in developing the next generation of architecture for high-performance computing
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Memory Systems Performance Architect
AI Memory Systems Performance Architect

CyberCoders • Santa Clara (CA)

On-site
USD 200,000 - 250,000
Relocation assistance and visa sponsor
Daily lunch stipend
Equity grant
+2
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Los Angeles (CA)

Hybrid
USD 266,000 - 445,000
Performance Modeling Architect
Performance Modeling Architect

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 300,000
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
Performance Architect
Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Performance Modeling Engineer
Performance Modeling Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Relocation assistance
Performance Modeling Engineer
Performance Modeling Engineer

The Consensus • San Jose (CA)

On-site
USD 150,000 - 230,000
Medical, dental, and vision coverage
Housing subsidy near Santana Row: $2k/
Relocation support to San Jose
+3
Senior Performance Modeling Architect, CPU Fabric and LLC
Senior Performance Modeling Architect, CPU Fabric and LLC

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior Memory System Architect
Senior Memory System Architect

NVIDIA • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Memory System Architect
Senior Memory System Architect

NVIDIA • Austin (TX)

On-site
USD 152,000 - 288,000
Equity
Benefits