Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Cerebras Systems is hiring a Staff GPU Inference SDET to lead the validation and reliability of our GPU inference stack and rack-scale systems. You will design automated test ecosystems for multi-node clusters, ensure numerical correctness, and drive production-grade reliability for high-speed inference workloads.
You will architect and implement validation pipelines spanning API services, serving workers, runtimes, and firmware, while measuring TTFT, ITL, throughput, and tail latency to guide
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
As a Staff GPU Inference SDET, you will be the founding quality, reliability, and validation lead for a new GPU Inference Development team. Working closely with engineering leads and cross-functional systems infrastructure teams, you will design, build, and scale the end-to-end release qualification and automated test ecosystem for our GPU inference stack and rack-scale accelerated compute fleets. In this high-impact role, you will be responsible for building automated test suites to validate multi-node GPU cluster bring‑up, verifying prefill worker optimizations, testing open-source and custom serving engines, and ensuring numerical correctness and performance stability under real‑world streaming workloads. You will be the primary technical anchor ensuring production-grade reliability, fault isolation, and peak inference performance across accelerated GPU infrastructure.