AI Cluster Network IO Engineer for Low-Latency Throughput

Cerebras

United States

Remote

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Cerebras Systems seeks a software developer for the Host and Network IO team to optimize bandwidth and latency on a distributed AI/HPC system. You will collaborate with AI application teams, cluster architects, and FPGA/ASIC groups to deploy robust IO solutions and reduce congestion.

The role demands proficiency in socket programming, RDMA Verbs, and performance tuning to improve real-world AI throughput. You will lead network debugging across large clusters and help shape architectural changes

Qualifications

  • Master's/PhD in Computer Science or Electrical Engineering with 1 year industry experience, or 3+ years industry experience.
  • Experience in large software environments.
  • Embedded systems, HW/SW co-design, and driver development.
  • Networking experience.

Responsibilities

  • Develop x86/ARM software to expose IO capabilities for AI/HPC teams.
  • Govern a generic IO API with multiple internal users.
  • Develop control and configuration subsystems interacting with Cerebras hardware.
  • Drive network performance debugging of large AI clusters.
  • Collect and analyze network statistics and traces to diagnose bottlenecks.
  • Develop telemetry tools to increase visibility into the network and IO datapath.
  • Optimize CPU/memory utilization with kernel bypass and zero-copy techniques.
  • Integrate advanced networking technologies and protocols.
  • Lead cross-functional technical projects across teams to improve IO solutions.
  • Foster clear communication across teams and stakeholders.

Skills

Advanced degree in CS/EE + 1yr or 3+yr
Experience in large software envs
Embedded systems HW/SW co-design
Network experience

Education

Master's/PhD in Computer Science or Electrical Engineering

Job description

Cerebras Systems seeks a software developer for the Host and Network IO team to optimize bandwidth and latency on a distributed AI/HPC system. You will collaborate with AI application teams, cluster architects, and FPGA/ASIC groups to deploy robust IO solutions and reduce congestion.

The role demands proficiency in socket programming, RDMA Verbs, and performance tuning to improve real-world AI throughput. You will lead network debugging across large clusters and help shape architectural changes

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems IO Engineer: Low-Latency Network & RDMA
AI Systems IO Engineer: Low-Latency Network & RDMA

Cerebras • Sunnyvale (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Software Engineer - Host and Network IO
Software Engineer - Host and Network IO

Cerebras • United States

Remote
USD 150,000 - 210,000
Senior FPGA Network IO Architect for Ultra‑Fast AI Clusters
Senior FPGA Network IO Architect for Ultra‑Fast AI Clusters

Cerebras • United States

Remote
USD 180,000 - 240,000
Software Engineer - Host and Network IO
Software Engineer - Host and Network IO

Cerebras • Sunnyvale (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Host & Network IO Software Engineer for AI/HPC
Host & Network IO Software Engineer for AI/HPC

Jobtailor • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Network Architect
Network Architect

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Network Architect
Network Architect

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
AI/HPC Performance Engineer: Scale Large AI Clusters
AI/HPC Performance Engineer: Scale Large AI Clusters

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Systems Engineer, High‑Speed AI Inference Networking
Systems Engineer, High‑Speed AI Inference Networking

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs / stock options
401(k) matching
+2
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits