Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Cerebras Systems seeks a software developer for the Host and Network IO team to optimize bandwidth and latency on a distributed AI/HPC system. You will collaborate with AI application teams, cluster architects, and FPGA/ASIC groups to deploy robust IO solutions and reduce congestion.
The role demands proficiency in socket programming, RDMA Verbs, and performance tuning to improve real-world AI throughput. You will lead network debugging across large clusters and help shape architectural changes
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
The Host and Network IO Team develops the full IO path implementation between a distributed system of server nodes, through the cluster, down to the custom RoCE network stack implemented in Cerebras' system, and over the proprietary IOs onto the WSE. As a software developer on the team, you will interface between AI application-level IO teams, cluster architecture teams, and FPGA/ASIC teams to develop solutions that optimize bandwidth and latency while minimizing congestion, pauses, pause spreading, unfairness, etc. Strong skills in socket programming will enable you to deploy robust management operations, while deftness in RDMA Verbs will enable you to optimize CPU resources and shape network traffic to deliver real world impact on AI performance metrics, as well as developing tools for gaining insight and visibility into network behavior. Meticulous analysis and rigour are key tenants of this role, harnessing that together with a deeply-understood mental model of the server, NIC, protocol, switch, and custom hardware behavior will enable you to lead network debug, optimize traffic patterns, and prescribe architectural changes.