Software Engineer - Host and Network IO

Cerebras

Toronto

On-site

CAD 120,000 - 180,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cerebras Systems is seeking a software developer for the Host and Network IO Team to optimize bandwidth and latency across a distributed AI hardware stack. You will interface with AI application IO teams, cluster architecture, and FPGA/ASIC groups to deploy robust, low-latency solutions on Cerebras hardware.

Responsibilities include developing x86/ARM software, governing IO APIs, and leading network performance debugging for large AI clusters, with emphasis on transparency and telemetry for

Qualifications

  • Master's/PhD in Computer Science or Electrical Engineering + 1 year industry experience, OR 3+ years industry experience.
  • Experience in large software environments.
  • Embedded systems, HW/SW co-design, and some driver development.
  • Network protocol familiarity (TCP, RoCE) and network debug tools such as Wireshark, or willingness to learn
  • Some network switch environment familiarity or willingness to learn (Arista, Juniper, etc.).
  • Detail-oriented but keen to learn the bigger picture and step out of comfort zone to embrace the unknown.

Responsibilities

  • Develop x86 & ARM software to expose next-generation hardware IO capabilities for AI/HPC application teams
  • Govern a generic IO API with multiple internal users.
  • Develop control and configuration subsystems directly interacting with Cerebras hardware
  • Drive network performance debug of large AI clusters
  • Gather and analyze network statistics and packet traces to root cause and alleviate bottlenecks and sub-optimalities.
  • Develop tools/telemetry for increasing visibility into the network and IO datapath.
  • Optimize cpu/mem utilization leveraging kernel bypass and zero-copy techniques
  • Integrate leading edge networking technologies and protocols
  • Lead cross-functional technical projects spanning multiple teams and integrating diverse software and hardware components to deliver an improved network IO solution.
  • Foster clear and effective communication across teams and stakeholders.

Skills

Embedded systems
HW/SW co-design
Driver development
TCP RoCE familiarity
Network debugging
Experience in large software Envs

Education

Master's/PhD in Computer Science or Electrical Engineering

Tools

Wireshark
Arista/Juniper familiarity

Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

About The Role

The Host and Network IO Team develops the full IO path implementation between a distributed system of server nodes, through the cluster, down to the custom RoCE network stack implemented in Cerebras' system, and over the proprietary IOs onto the WSE. As a software developer on the team, you will interface between AI application-level IO teams, cluster architecture teams, and FPGA/ASIC teams to develop solutions that optimize bandwidth and latency while minimizing congestion, pauses, pause spreading, unfairness, etc. Strong skills in socket programming will enable you to deploy robust management operations, while deftness in RDMA Verbs will enable you to optimize CPU resources and shape network traffic to deliver real world impact on AI performance metrics, as well as developing tools for gaining insight and visibility into network behavior. Meticulous analysis and rigour are key tenants of this role, harnessing that together with a deeply-understood mental model of the server, NIC, protocol, switch, and custom hardware behavior will enable you to lead network debug, optimize traffic patterns, and prescribe architectural changes.

Responsibilities
  • Develop x86 & ARM software to expose next-generation hardware IO capabilities for AI/HPC application teams
  • Govern a generic IO API with multiple internal users.
  • Develop control and configuration subsystems directly interacting with Cerebras hardware
  • Drive network performance debug of large AI clusters
  • Gather and analyze network statistics and packet traces to root cause and alleviate bottlenecks and sub-optimalities.
  • Develop tools/telemetry for increasing visibility into the network and IO datapath.
  • Optimize cpu/mem utilization leveraging kernel bypass and zero-copy techniques
  • Integrate leading edge networking technologies and protocols
  • Lead cross-functional technical projects spanning multiple teams and integrating diverse software and hardware components to deliver an improved network IO solution.
  • Foster clear and effective communication across teams and stakeholders.
Skills & Qualifications
  • Master's/PhD in Computer Science or Electrical Engineering + 1 year industry experience, OR 3+ years industry experience.
  • Experience in large software environments.
  • Embedded systems, HW/SW co-design, and some driver development.
  • Network protocol familiarity (TCP, RoCE) and network debug tools such as Wireshark, or willingness to learn
  • Some network switch environment familiarity or willingness to learn (Arista, Juniper, etc.).
  • Detail-oriented but keen to learn the bigger picture and step out of comfort zone to embrace the unknown.
Why Join Cerebras

People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

FPGA Engineer
FPGA Engineer

Foundation Capital • Toronto

On-site
CAD 120,000 - 180,000
FPGA Engineer
FPGA Engineer

Cerebras Systems • Lower Sackville

On-site
CAD 100,000 - 170,000
FPGA Engineer
FPGA Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 160,000
Distributed Software Engineer
Distributed Software Engineer

Engg • Toronto

On-site
CAD 140,000 - 210,000
Software Engineer - Tools & Infrastructure / DevOps
Software Engineer - Tools & Infrastructure / DevOps

Foundation Capital • Toronto

On-site
CAD 110,000 - 170,000
Software Engineer - Tools & Infrastructure / DevOps
Software Engineer - Tools & Infrastructure / DevOps

Cerebras • Toronto

On-site
CAD 120,000 - 180,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Foundation Capital • Toronto

On-site
CAD 120,000 - 180,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Toronto

On-site
CAD 110,000 - 160,000
AI Inference Core - Senior SW Engineer for Platform & DevOps
AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras • Toronto

On-site
CAD 120,000 - 180,000
Software Engineer, Inference Platform
Software Engineer, Inference Platform

Foundation Capital • Toronto

On-site
CAD 256,000 - 370,000