Senior AI Systems Engineer - Kernel & GPU Networking

Cisco

Milpitas (CA)

On-site

USD 195,000 - 218,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cisco is seeking a Senior Software Engineer to join the CVIS team in Milpitas, CA, responsible for debugging, profiling and optimizing the AI cluster networking and GPU performance stack from boot-time through the Linux kernel, RDMA, NCCL and application-level communications.

You will investigate kernel-level failures, boot-time anomalies, PCIe issues, and driver behavior, diagnose and tune RDMA/RoCEv2, NCCL, SmartNIC/DPU, and NIC queues, and read traces, logs and observability data to identify

Qualifications

  • 4+ years of experience working with Linux kernel, networking, GPU, and distributed-systems debugging.
  • Proficient in C/C++/Python and driver development (user-mode/kernel-mode).
  • Experience with Linux networking stack, RDMA/RoCEv2, NCCL, PCIe and performance profiling.

Responsibilities

  • Investigate kernel-level failures, boot-time anomalies, PCIe link issues, driver behavior, and network latency.
  • Diagnose and tune RDMA/RoCEv2, NCCL, SmartNIC/DPU, NIC queues, congestion, throughput, and collective communications.
  • Read and correlate unified hardware traces across host CPUs, PCIe switches, SmartNICs, and GPU compute streams.
  • Capture and analyze Wireshark packets, system logs, NVIDIA support logs, traces, profiles, and observability data.
  • Partner with hardware and software teams to eliminate end-to-end bottlenecks.

Skills

Linux kernel
Networking
GPU
Distributed systems debugging
C/C++/Python
RDMA
NCCL
PCIe
Performance profiling

Education

Bachelor's degree in Computer Science or related field

Tools

Nsight profiling tools

Job description

Cisco is seeking a Senior Software Engineer to join the CVIS team in Milpitas, CA, responsible for debugging, profiling and optimizing the AI cluster networking and GPU performance stack from boot-time through the Linux kernel, RDMA, NCCL and application-level communications.

You will investigate kernel-level failures, boot-time anomalies, PCIe issues, and driver behavior, diagnose and tune RDMA/RoCEv2, NCCL, SmartNIC/DPU, and NIC queues, and read traces, logs and observability data to identify

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Cluster Performance Engineer
Senior AI Cluster Performance Engineer

020 Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 168,000 - 245,000
Senior AI Cluster Validation & Benchmarking Engineer
Senior AI Cluster Validation & Benchmarking Engineer

Cisco Systems, Inc • Milpitas (CA)

On-site
USD 168,000 - 245,000
Senior AI Cluster Validation & Release Engineer
Senior AI Cluster Validation & Release Engineer

020 Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 180,000 - 260,000
Medical benefits
Dental benefits
Vision benefits
+6
Senior AI Cluster Validation & Performance Engineer
Senior AI Cluster Validation & Performance Engineer

Cisco • Milpitas (CA)

On-site
USD 168,000 - 245,000
Medical, dental, vision insurance
401(k) plan with Cisco matching
Paid parental leave
+3
Senior AI Cluster Network Solutions Engineer
Senior AI Cluster Network Solutions Engineer

NVIDIA • Durham (NC)

On-site
USD 168,000 - 322,000
Equity
Benefits package
Senior AI Cluster Network Engineer
Senior AI Cluster Network Engineer

NVIDIA • Seattle (WA)

On-site
USD 168,000 - 322,000
Senior AI Networking Solutions Engineer
Senior AI Networking Solutions Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 322,000
Equity
Benefits
Senior AI Networking Performance Architect — Equity
Senior AI Networking Performance Architect — Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior Solutions Engineer - AI Networking & HPC (Equity)
Senior Solutions Engineer - AI Networking & HPC (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 322,000
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 431,250