Senior AI Systems Performance Modeling Architect

Oho Group

San Francisco (CA)

On-site

USD 180,000 - 300,000

Full time

37 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Oho Group in San Francisco, CA is seeking a Principal Performance Modeling Architect for AI systems to lead architecture-informed performance modeling before hardware realization. You will build analytical and simulation models to guide decisions across compute, memory, and interconnect, influencing silicon and software direction.

The role emphasizes strong Python/C++ skills, deep architectural knowledge, and collaboration with hardware, compiler, and runtime teams to translate model results

Qualifications

  • Strong Python and/or C++ development experience.
  • Experience building performance models for architectures.
  • Deep knowledge of computer architecture and microarchitecture.
  • Understanding compute pipelines, memory systems and data movement.
  • Ability to translate model results into architectural decisions.
  • Experience with GPUs, AI accelerators, LLM inference is a plus.

Responsibilities

  • Build analytical and simulation-based performance models.
  • Model processors, accelerators, memory hierarchies and interconnects.
  • Analyze transformer inference across single-device and multi-accelerator systems.
  • Evaluate latency, throughput, utilization, energy and cost trade-offs.
  • Characterize workload behavior using traces and benchmarks.
  • Identify architectural bottlenecks and propose measurable improvements.
  • Partner with hardware, compiler, runtime and inference teams.
  • Correlate models against simulation, emulation or silicon as the platform matures.

Skills

Python
C++
Performance modeling

Job description

Oho Group in San Francisco, CA is seeking a Principal Performance Modeling Architect for AI systems to lead architecture-informed performance modeling before hardware realization. You will build analytical and simulation models to guide decisions across compute, memory, and interconnect, influencing silicon and software direction.

The role emphasizes strong Python/C++ skills, deep architectural knowledge, and collaboration with hardware, compiler, and runtime teams to translate model results

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Modeling Architect
Performance Modeling Architect

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 300,000
Performance Modeling Engineer - AI Systems & Infrastructure
Performance Modeling Engineer - AI Systems & Infrastructure

OpenAI • Seattle (WA)

Hybrid
USD 293,000 - 385,000
Relocation assistance
Hybrid work model
AI Systems Performance Modeling Architect
AI Systems Performance Modeling Architect

Velaura • Santa Clara (CA)

On-site
USD 200,000 - 500,000
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
Performance Modeling Engineer
Performance Modeling Engineer

OpenAI • Los Angeles (CA)

Hybrid
USD 100,000 - 150,000
Relocation assistance
Hybrid work model
AI Compute Performance Modeling Architect
AI Compute Performance Modeling Architect

Velaura AI • North Carolina

On-site
USD 150,000 - 185,000
Performance Modeling Engineer
Performance Modeling Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Relocation assistance
Performance Modeling Lead
Performance Modeling Lead

OpenAI • San Francisco (CA)

Hybrid
USD 342,000 - 555,000
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Los Angeles (CA)

Hybrid
USD 266,000 - 445,000
Lead Architect, AI Performance & Modeling (Hybrid)
Lead Architect, AI Performance & Modeling (Hybrid)

d-Matrix inc. • Santa Clara (CA)

Hybrid
USD 150,000 - 200,000