Stand out for this role — generate a tailored resume and cover letter in about a minute.
Oho Group in San Francisco seeks a senior performance-modeling engineer to build analytical and simulation-based models that guide architectural decisions for AI workloads before hardware exists.
Your work will span processors, accelerators, memory hierarchies and interconnects, analyzing transformer inference across single-device and multi-accelerator systems, and evaluating latency, throughput, energy and cost trade-offs.
We’re working with a semiconductor company designing a new programmable compute architecture for demanding AI workloads.
They’re looking for a senior performance-modeling engineer to create the tools that guide architectural decisions before hardware exists. Your work will determine where performance is gained or lost across compute, memory, interconnect and complete AI systems.
Experience with GPUs, AI accelerators, LLM inference, cluster modelling, roofline analysis, NoCs or cost-per-token analysis would be especially relevant.
This is an architecture-shaping role where performance models influence both the silicon and the software designed around it.