Performance Modeling Engineer

Oho Group

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

10 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oho Group in San Francisco seeks a senior performance-modeling engineer to build analytical and simulation-based models that guide architectural decisions for AI workloads before hardware exists.

Your work will span processors, accelerators, memory hierarchies and interconnects, analyzing transformer inference across single-device and multi-accelerator systems, and evaluating latency, throughput, energy and cost trade-offs.

Qualifications

  • Strong Python and/or C++ development.
  • Experience building performance models.
  • Deep computer-architecture and microarchitecture knowledge.
  • Understanding of compute pipelines, memory systems and data movement.
  • Ability to convert model results into concrete architecture decisions.

Responsibilities

  • Build analytical and simulation-based performance models.
  • Model processors, accelerators, memory hierarchies and interconnects.
  • Analyze transformer inference across single-device and multi-accelerator systems.
  • Evaluate latency, throughput, utilization, energy and cost trade-offs.
  • Characterize workload behavior using traces and representative benchmarks.
  • Identify architectural bottlenecks and propose measurable improvements.
  • Partner with hardware, compiler, runtime and inference teams.
  • Correlate models against simulation, emulation or silicon as the platform matures.

Skills

Python
C++
Performance modeling
Computer architecture

Job description

We’re working with a semiconductor company designing a new programmable compute architecture for demanding AI workloads.

They’re looking for a senior performance-modeling engineer to create the tools that guide architectural decisions before hardware exists. Your work will determine where performance is gained or lost across compute, memory, interconnect and complete AI systems.

What you’ll work on
  • Build analytical and simulation-based performance models
  • Model processors, accelerators, memory hierarchies and interconnects
  • Analyze transformer inference across single-device and multi-accelerator systems
  • Evaluate latency, throughput, utilization, energy and cost trade-offs
  • Characterize workload behavior using traces and representative benchmarks
  • Identify architectural bottlenecks and propose measurable improvements
  • Partner with hardware, compiler, runtime and inference teams
  • Correlate models against simulation, emulation or silicon as the platform matures
What we’re looking for
  • Strong Python and/or C++ development
  • Experience building performance models
  • Deep computer-architecture and microarchitecture knowledge
  • Understanding of compute pipelines, memory systems and data movement
  • Ability to convert model results into concrete architecture decisions

Experience with GPUs, AI accelerators, LLM inference, cluster modelling, roofline analysis, NoCs or cost-per-token analysis would be especially relevant.

This is an architecture-shaping role where performance models influence both the silicon and the software designed around it.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

Hybrid
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
Performance Architect
Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Seattle (WA)

On-site
USD 342,000 - 555,000
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Los Angeles (CA)

Hybrid
USD 266,000 - 445,000
Performance Architect - AI Hardware
Performance Architect - AI Hardware

TEEMA • United States

Remote
USD 140,000 - 210,000
Member of Technical Staff, Performance Modeling
Member of Technical Staff, Performance Modeling

Socket.dev • Santa Clara (CA)

On-site
USD 170,000 - 250,000
Health, Dental, Vision coverage
401(k) match
Equity grant
+2
Member of Technical Staff, Performance Modeling
Member of Technical Staff, Performance Modeling

Netpreme • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Health, Dental, and Vision coverage
Early-stage equity grant
401(k) match
+2
Member of Technical Staff, Performance Modeling
Member of Technical Staff, Performance Modeling

Netpreme • Boston (MA)

On-site
USD 150,000 - 230,000
Health coverage
Dental coverage
Vision coverage
+12
Performance Modeling Lead
Performance Modeling Lead

OpenAI • San Francisco (CA)

Hybrid
USD 342,000 - 555,000
Performance Modeling Engineer
Performance Modeling Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Relocation assistance