Senior AI Cluster Performance Engineer

020 Cisco Systems, Inc.

Milpitas (CA)

On-site

USD 168,000 - 245,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Cisco Systems, Inc. is seeking a Senior Software Engineer to debug, profile, and optimize the AI cluster networking and GPU performance stack across boot-time to kernel level, SmartNIC/DPU, and NCCL-based communications.

You will investigate kernel-level failures, tune RDMA and NCCL, and read hardware traces to identify end-to-end bottlenecks while collaborating with hardware and software teams to drive performance improvements.

Qualifications

  • Bachelor's degree + 7 years of related experience, or Master's degree + 4 years, or PhD + 1 year, or equivalent related work experience.
  • 4+ years of experience with Linux kernel, networking, GPU, and distributed-systems debugging.
  • Experience writing code in C/C++/Python and experience with user-mode and kernel-mode drivers.
  • Experience with Linux network stack, RDMA/RoCEv2, NCCL, PCIe, and performance profiling.
  • Experience reading packet captures, traces, counters, logs, and reproducible benchmarks.

Responsibilities

  • Investigate kernel-level failures, boot-time anomalies, PCIe link issues, driver behavior, and network latency.
  • Diagnose and tune RDMA/RoCEv2, NCCL, SmartNIC/DPU, NIC queues, congestion, throughput, and collective communications.
  • Read and correlate unified hardware traces across host CPUs, PCIe switches, SmartNICs, and GPU compute streams.
  • Capture and analyze Wireshark packets, system logs, NVIDIA support logs, traces, profiles, and observability data.
  • Partner with hardware and software teams to eliminate end-to-end bottlenecks.

Skills

C/C++/Python
Linux kernel
Networking
Distributed debugging
RDMA/RoCEv2
NCCL
PCIe

Education

Bachelor's degree
Master's degree
PhD

Tools

Nsight profiling tools
Kernel tracing
Observability platforms
Wireshark

Job description

Cisco Systems, Inc. is seeking a Senior Software Engineer to debug, profile, and optimize the AI cluster networking and GPU performance stack across boot-time to kernel level, SmartNIC/DPU, and NCCL-based communications.

You will investigate kernel-level failures, tune RDMA and NCCL, and read hardware traces to identify end-to-end bottlenecks while collaborating with hardware and software teams to drive performance improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer - Kernel & GPU Networking
Senior AI Systems Engineer - Kernel & GPU Networking

Cisco • Milpitas (CA)

On-site
USD 195,000 - 218,000
Senior AI Cluster Validation & Performance Engineer
Senior AI Cluster Validation & Performance Engineer

Cisco • Milpitas (CA)

On-site
USD 168,000 - 245,000
Medical, dental, vision insurance
401(k) plan with Cisco matching
Paid parental leave
+3
Senior AI Cluster Validation & Benchmarking Engineer
Senior AI Cluster Validation & Benchmarking Engineer

Cisco Systems, Inc • Milpitas (CA)

On-site
USD 168,000 - 245,000
AI Compute Performance Engineer
AI Compute Performance Engineer

Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 155,000 - 223,000
Medical, dental, vision insurance
401(k) with Cisco matching
Paid parental leave
+2
Senior AI Networking Performance Architect — Equity
Senior AI Networking Performance Architect — Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior AI Cluster Network Solutions Engineer
Senior AI Cluster Network Solutions Engineer

NVIDIA • Durham (NC)

On-site
USD 168,000 - 322,000
Equity
Benefits package
Senior AI Cluster Validation & Release Engineer
Senior AI Cluster Validation & Release Engineer

020 Cisco Systems, Inc. • Milpitas (CA)

On-site
USD 180,000 - 260,000
Medical benefits
Dental benefits
Vision benefits
+6
Senior Networking Solutions Engineer for AI Clusters
Senior Networking Solutions Engineer for AI Clusters

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 168,000 - 322,000
Equity
Benefits package
Senior AI Cluster Network Engineer
Senior AI Cluster Network Engineer

NVIDIA • Seattle (WA)

On-site
USD 168,000 - 322,000
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 431,250