Get more replies from employers
Send a job-specific resume in minutes.
SambaNova Systems is seeking an Architect on the Inference Systems Performance team to own end-to-end performance for large-scale LLM inference, from tokenization to deployment, across hardware and software components. You will drive reproducible workload capture, benchmarking, modeling, and simulation to inform capacity planning against customer SLOs.
Lead the performance-modeling effort, mentor senior engineers, and collaborate with model optimization, systems, and product teams to balance
The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale. SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.
As an Architect on the Inference Systems Performance team, you'll own the discipline of end-to-end performance for large-scale LLM inference at SambaNova, from how a request moves through tokenization, prefill, decode, and the fabric between them, to how an entire deployment is sized against customer SLOs. Inference-systems performance is a nascent field; the results of design choices are being discovered daily rather than inherited from a mature craft, and this role exists to bring rigor to that frontier. The work spans two coupled pillars. The first is reproducible workload capture and benchmarking -- building faithful, replayable representations of real and increasingly agentic traffic, so that what we measure reflects production rather than an artifact of a naive load script. The second is performance modeling and simulation - analytic and simulation models that turn measurement into a "what-if" capability, letting us reason about configurations and hardware that do not exist yet. Together these feed both today's serving optimization and the next generation of system planning. The technical frontier you'll help define is heterogeneous, disaggregated inference - GPU on prefill, the RDU on decode - which explores hard problems across networking, storage, prompt caching, and tail-latency-bound data movement. You will be the go-to person for inference-systems performance across SambaNova, and a resource the entire organization relies on to answer "how fast can this go, and what will it take."
Base Salary Range: Base Pay Range $245,000 - $325,000 USD
SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.