A complete application in a minute — tailored resume and cover letter, ready to send.
Renice AI is seeking a Performance Lead to drive architecture decisions for AI inference workloads across the full system stack in Mountain View, CA. You will build modeling frameworks, analyze tradeoffs across compute, memory, networking, and storage, and translate results into actionable guidance for internal teams and hardware partners.
This role leads a small team of engineers, collaborates with ML, systems, and hardware groups, and shapes reference designs and long-term infrastructure
Renice AI develops system and infrastructure solutions designed for the unique demands of advanced AI inference workloads. We work closely with external research, software, and hardware partners to shape the next generation of AI systems, from silicon to model weights through full-scale deployments.
This role focuses on understanding and optimizing performance across the full system stack, ensuring that architectural decisions are grounded in rigorous, quantitative analysis of real-world workloads.
We are seeking a Performance Lead to answer forward-looking architectural questions across AI infrastructure systems.
You will develop modeling frameworks and methodologies to evaluate system-level tradeoffs and guide key design decisions. Your work will directly influence reference architectures, vendor designs, and long-term infrastructure strategy.
This role sits at the intersection of AI workloads, system architecture, and quantitative modeling. It requires strong technical judgment, ownership, and the ability to translate complex analysis into clear, actionable guidance.
Renice AI builds the integrated stack for AI inference: the model, the software and the hardware designed as one system, rather than one layer at a time.
Nobody designed the AI stack as a whole. Models, software, chips, memory and networking are each owned by a different part of the industry, and each locked in choices that were rational at the time but were never made for serving. Added up, those choices cap how many tokens the world can make. We redesign across them so the machines the world already has produce more tokens in the datacenter and on the desktop.