A complete application in a minute — tailored resume and cover letter, ready to send.
Cerebras Systems seeks an experienced FPGA developer to own the in-chassis IO subsystem, including RoCE interfaces, a large programmable switching fabric, and the Cerebras WSE communication protocol. You will optimize bandwidth and latency for multi-node AI clusters and deliver production‑ready solutions.
You will collaborate with AI application IO teams, cluster architecture, and embedded software groups across hardware and software boundaries to enable scalable, high‑speed inference and
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
The Host and Network IO Team develops the full IO path implementation between a distributed system of server nodes, through the cluster, down to the custom RoCE network stack implemented in Cerebras' system, and over the proprietary IOs onto the WSE. As an FPGA developer on the team, you will own the in-chassis IO subsystem consisting of i) several cluster-facing RoCE v2 network interfaces via a custom implementation of the RDMA protocol; ii) a large programmable switching fabric; and iii) Serial IO communication with the Cerebras WSE via a proprietary protocol. You will interface between AI application-level IO teams, cluster architecture teams, and embedded software teams to develop solutions that optimize bandwidth and latency while minimizing congestion, pauses, pause spreading, unfairness, etc. The scope of work spans multiple generations of hardware products from improvements to presently deployed hardware, implementation of upcoming systems, and design/architecting of future next-gen architectures.
Lead full chassis-to-wafer IO architecture and design
Improve current RTL, implement next-gen FPGA design, define future IO architectures
Produce production-ready bitstreams for deployment into large clusters running customer inference services
Work with DV team to prevent bug slips and simplify debug
Optimize bandwidth/latency over FPGA datapath and all IO interfaces
Interface with board team facilitating and executing board bringup
Drive network performance debug of large AI clusters
Integrate leading edge networking technologies and protocols
Lead cross-functional technical projects spanning multiple teams and integrating diverse software and hardware components to deliver an improved network IO solution.
Foster clear and effective communication across teams and stakeholders.
Why Join Cerebras
People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:
Find out more about what it's like to work at Cerebras here!
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.