Performance Modeling Architect

Oho Group

San Francisco (CA)

On-site

USD 180,000 - 300,000

Full time

27 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Oho Group in San Francisco, CA is seeking a Principal Performance Modeling Architect for AI systems to lead architecture-informed performance modeling before hardware realization. You will build analytical and simulation models to guide decisions across compute, memory, and interconnect, influencing silicon and software direction.

The role emphasizes strong Python/C++ skills, deep architectural knowledge, and collaboration with hardware, compiler, and runtime teams to translate model results

Qualifications

  • Strong Python and/or C++ development experience.
  • Experience building performance models for architectures.
  • Deep knowledge of computer architecture and microarchitecture.
  • Understanding compute pipelines, memory systems and data movement.
  • Ability to translate model results into architectural decisions.
  • Experience with GPUs, AI accelerators, LLM inference is a plus.

Responsibilities

  • Build analytical and simulation-based performance models.
  • Model processors, accelerators, memory hierarchies and interconnects.
  • Analyze transformer inference across single-device and multi-accelerator systems.
  • Evaluate latency, throughput, utilization, energy and cost trade-offs.
  • Characterize workload behavior using traces and benchmarks.
  • Identify architectural bottlenecks and propose measurable improvements.
  • Partner with hardware, compiler, runtime and inference teams.
  • Correlate models against simulation, emulation or silicon as the platform matures.

Skills

Python
C++
Performance modeling

Job description

Principal Performance Modeling Architect — AI Systems

We’re working with a semiconductor company designing a new programmable compute architecture for demanding AI workloads.

They’re looking for a senior performance-modeling engineer to create the tools that guide architectural decisions before hardware exists. Your work will determine where performance is gained or lost across compute, memory, interconnect and complete AI systems.

What you’ll work on
  • Build analytical and simulation-based performance models
  • Model processors, accelerators, memory hierarchies and interconnects
  • Analyze transformer inference across single-device and multi-accelerator systems
  • Evaluate latency, throughput, utilization, energy and cost trade-offs
  • Characterize workload behavior using traces and representative benchmarks
  • Identify architectural bottlenecks and propose measurable improvements
  • Partner with hardware, compiler, runtime and inference teams
  • Correlate models against simulation, emulation or silicon as the platform matures
What we’re looking for
  • Strong Python and/or C++ development
  • Experience building performance models
  • Deep computer-architecture and microarchitecture knowledge
  • Understanding of compute pipelines, memory systems and data movement
  • Ability to convert model results into concrete architecture decisions

Experience with GPUs, AI accelerators, LLM inference, cluster modelling, roofline analysis, NoCs or cost-per-token analysis would be especially relevant.

This is an architecture-shaping role where performance models influence both the silicon and the software designed around it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
Senior AI Systems Performance Modeling Architect
Senior AI Systems Performance Modeling Architect

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 300,000
Performance Architect
Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Performance Architect - AI Hardware
Performance Architect - AI Hardware

TEEMA • United States

Remote
USD 140,000 - 210,000
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Los Angeles (CA)

Hybrid
USD 266,000 - 445,000
AI Performance Modeling Architect for CPU/SoC Design
AI Performance Modeling Architect for CPU/SoC Design

Actalent • Raleigh (NC)

Hybrid
USD 200,000 - 500,000
Hybrid work in Raleigh
Performance Modeling Engineer
Performance Modeling Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Relocation assistance
Performance Modeling Engineer - AI Systems & Infrastructure
Performance Modeling Engineer - AI Systems & Infrastructure

OpenAI • Seattle (WA)

Hybrid
USD 293,000 - 385,000
Relocation assistance
Hybrid work model
Performance Modeling Lead
Performance Modeling Lead

OpenAI • San Francisco (CA)

Hybrid
USD 342,000 - 555,000
Performance Modeling Architect - AI Memory Systems
Performance Modeling Architect - AI Memory Systems

CyberCoders • Santa Clara (CA)

On-site
USD 200,000 - 250,000
Relocation assistance and visa sponsor
Daily lunch stipend
Equity grant
+2