Senior Principal Engineer - AI Networking

Oracle Corporation

Seattle (WA)

On-site

USD 180,000 - 260,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oracle Corporation in Seattle seeks a senior architect to lead high-performance networking software for large-scale AI and HPC environments, focusing on RDMA-based communication across GPU clusters.

You will mentor engineers, shape architecture, and collaborate across hardware, networking, and AI platform teams to deliver scalable infrastructure. Experience with NCCL, UCX, and GPU networking preferred.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or related field; advanced degree preferred.
  • 10+ years of software engineering experience building distributed systems, networking software, or infrastructure platforms.
  • Deep expertise in RDMA technologies including RoCE, InfiniBand, or equivalent high-performance networking technologies.
  • Strong experience developing networking software in C/C++.
  • Experience designing and optimizing distributed communication frameworks and transport protocols.
  • Solid understanding of operating systems, networking stacks, memory management, and performance optimization.
  • Experience troubleshooting and optimizing large-scale production systems.
  • Demonstrated technical leadership driving architecture and execution across multiple teams.
  • Strong knowledge of Linux systems and low-level systems programming.

Responsibilities

  • Architect and develop high-performance networking software for large-scale AI and HPC environments.
  • Design and implement RDMA-based services and infrastructure enabling low-latency, high-throughput communication across GPU clusters.
  • Drive evolution of collective communication frameworks and transport layers for distributed AI workloads.
  • Develop congestion management, traffic engineering, load balancing, and resiliency for large-scale RDMA networks.
  • Optimize end-to-end communication performance across networking, GPU, and software stacks.
  • Collaborate with hardware, networking, distributed systems, and AI platform teams to deliver scalable infrastructure.

Skills

Distributed systems
RDMA
Performance tuning
Leadership
Mentoring
Communication

Education

Bachelor's degree in CS/EE
Advanced degree preferred

Tools

C/C++
NCCL
RDMA
GPUDirect RDMA

Job description

You will work at the intersection of distributed systems, networking, and AI infrastructure, driving architecture, design, implementation, and performance optimization across software components that support thousands of GPUs and high-bandwidth network fabrics. The ideal candidate combines deep expertise in RDMA and distributed communication systems with a strong track record of delivering production-grade infrastructure at scale.

As a technical leader, you will influence architecture across multiple teams, mentor senior engineers, and help shape the roadmap for Oracle's AI networking platform.

What You'll Bring

  • Ability to solve highly complex technical challenges spanning networking, distributed systems, and AI infrastructure.
  • Strong system design skills with a focus on scalability, performance, and reliability.
  • A data-driven approach to performance analysis and optimization.
  • Excellent communication and collaboration skills across engineering organizations.
  • Passion for building foundational technologies that enable the next generation of AI workloads.
Internal Responsibilities

Key Responsibilities

  • Architect and develop high-performance networking software for large-scale AI and HPC environments.
  • Design and implement RDMA-based services and infrastructure that enable low-latency, high-throughput communication across GPU clusters.
  • Drive the evolution of collective communication frameworks and transport layers used by distributed AI training and inference workloads.
  • Develop congestion management, traffic engineering, load balancing, and resiliency mechanisms for large-scale RDMA networks.
  • Optimize end-to-end communication performance across networking, GPU, and software stacks.
  • Collaborate with hardware, networking, distributed systems, and AI platform teams to deliver scalable infrastructure solutions.
  • Lead performance analysis, bottleneck identification, and system-wide optimization efforts.
  • Define architecture and technical direction for networking platforms supporting next-generation AI workloads.
  • Build observability, monitoring, telemetry, and debugging capabilities for large-scale distributed systems.
  • Drive reliability, fault tolerance, and recovery mechanisms for mission-critical AI infrastructure.
  • Mentor engineers across the organization and provide technical leadership on complex cross-functional initiatives.
  • Influence engineering best practices, architecture reviews, and long-term technology strategy.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or related field; advanced degree preferred.
  • 10+ years of software engineering experience building distributed systems, networking software, or infrastructure platforms.
  • Deep expertise in RDMA technologies including RoCE, InfiniBand, or equivalent high-performance networking technologies.
  • Strong experience developing networking software in C/C++.
  • Experience designing and optimizing distributed communication frameworks and transport protocols.
  • Solid understanding of operating systems, networking stacks, memory management, and performance optimization.
  • Experience troubleshooting and optimizing large-scale production systems.
  • Demonstrated technical leadership driving architecture and execution across multiple teams.
  • Strong knowledge of Linux systems and low-level systems programming.

Preferred Qualifications

  • Experience with collective communication libraries such as NCCL, RCCL, MPI, UCC, UCX, XCCL, or similar technologies.
  • Experience building AI infrastructure supporting distributed training and inference workloads.
  • Expertise in GPU networking technologies including GPUDirect RDMA and GPU-aware communication stacks.
  • Experience with congestion management, adaptive routing, traffic shaping, and network resiliency mechanisms.
  • Familiarity with large-scale GPU clusters consisting of hundreds to thousands of accelerators.
  • Experience developing services and platforms operating directly over RDMA transports.
  • Knowledge of distributed training frameworks such as PyTorch, DeepSpeed, Megatron-LM, TensorFlow, or JAX.
  • Experience with cloud infrastructure and large-scale production service deployment.
  • Familiarity with Kubernetes, containerized environments, and cloud-native infrastructure.
  • Experience leading architecture for highly available and performance-critical systems.
External Responsibilities

Key Responsibilities

  • Architect and develop high-performance networking software for large-scale AI and HPC environments.
  • Design and implement RDMA-based services and infrastructure that enable low-latency, high-throughput communication across GPU clusters.
  • Drive the evolution of collective communication frameworks and transport layers used by distributed AI training and inference workloads.
  • Develop congestion management, traffic engineering, load balancing, and resiliency mechanisms for large-scale RDMA networks.
  • Optimize end-to-end communication performance across networking, GPU, and software stacks.
  • Collaborate with hardware, networking, distributed systems, and AI platform teams to deliver scalable infrastructure solutions.
  • Lead performance analysis, bottleneck identification, and system-wide optimization efforts.
  • Define architecture and technical direction for networking platforms supporting next-generation AI workloads.
  • Build observability, monitoring, telemetry, and debugging capabilities for large-scale distributed systems.
  • Drive reliability, fault tolerance, and recovery mechanisms for mission-critical AI infrastructure.
  • Mentor engineers across the organization and provide technical leadership on complex cross-functional initiatives.
  • Influence engineering best practices, architecture reviews, and long-term technology strategy.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or related field; advanced degree preferred.
  • 10+ years of software engineering experience building distributed systems, networking software, or infrastructure platforms.
  • Deep expertise in RDMA technologies including RoCE, InfiniBand, or equivalent high-performance networking technologies.
  • Strong experience developing networking software in C/C++.
  • Experience designing and optimizing distributed communication frameworks and transport protocols.
  • Solid understanding of operating systems, networking stacks, memory management, and performance optimization.
  • Experience troubleshooting and optimizing large-scale production systems.
  • Demonstrated technical leadership driving architecture and execution across multiple teams.
  • Strong knowledge of Linux systems and low-level systems programming.

Preferred Qualifications

  • Experience with collective communication libraries such as NCCL, RCCL, MPI, UCC, UCX, XCCL, or similar technologies.
  • Experience building AI infrastructure supporting distributed training and inference workloads.
  • Expertise in GPU networking technologies including GPUDirect RDMA and GPU-aware communication stacks.
  • Experience with congestion management, adaptive routing, traffic shaping, and network resiliency mechanisms.
  • Familiarity with large-scale GPU clusters consisting of hundreds to thousands of accelerators.
  • Experience developing services and platforms operating directly over RDMA transports.
  • Knowledge of distributed training frameworks such as PyTorch, DeepSpeed, Megatron-LM, TensorFlow, or JAX.
  • Experience with cloud infrastructure and large-scale production service deployment.
  • Familiarity with Kubernetes, containerized environments, and cloud-native infrastructure.
  • Experience leading architecture for highly available and performance-critical systems.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Oracle Corporation • Nashville (TN)

On-site
USD 180,000 - 240,000
Senior Principal Engineer - AI Networking
Senior Principal Engineer - AI Networking

Ll Oefentherapie • Seattle (WA)

On-site
USD 120,000 - 160,000
Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Oracle Corporation • Seattle (WA)

On-site
USD 150,000 - 210,000
Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 120,000 - 190,000
Lead Principal Software Engineer, Core Infrastructure
Lead Principal Software Engineer, Core Infrastructure

Oracle Corporation • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff HPC Network Architect
Staff HPC Network Architect

Lambda • United States

Hybrid
USD 180,000 - 280,000
Senior Core Infrastructure Engineer, AI Infrastructure
Senior Core Infrastructure Engineer, AI Infrastructure

Ll Oefentherapie • Nashville (TN)

On-site
USD 130,000 - 190,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Senior Manager - AI Network Engineering
Senior Manager - AI Network Engineering

Ll Oefentherapie • Nashville (TN)

On-site
USD 170,000 - 355,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off and holidays