Get more replies from employers
Send a job-specific resume in minutes.
SambaNova Systems is seeking an Architect for the Inference Systems Performance team to own end-to-end performance of large-scale LLM inference. You will study how requests move from tokenization to decode and how a deployment is sized to meet customer SLOs.
This role spans workload capture, benchmarking, and performance modeling to guide capacity planning and system design. You'll define strategies, build benchmarking capabilities, and develop tooling to profile a distributed inference
SambaNova Systems is seeking an Architect for the Inference Systems Performance team to own end-to-end performance of large-scale LLM inference. You will study how requests move from tokenization to decode and how a deployment is sized to meet customer SLOs.
This role spans workload capture, benchmarking, and performance modeling to guide capacity planning and system design. You'll define strategies, build benchmarking capabilities, and develop tooling to profile a distributed inference