Senior Network Reliability Engineer - DGX Cloud

NVIDIA Corporation

Santa Clara, Northern (CA, KY)

Hybrid

USD 136,000 - 265,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA Corporation seeks a Senior Network Reliability Engineer to support and maintain cloud and datacenter networks across NVIDIA’s software stack. You will remediate alerts, triage incidents, and engage with vendors for hardware/software issues.

The role requires hands-on troubleshooting, extensive TCP/IP and routing protocol knowledge, and experience with multi-vendor networks in cloud/data center contexts.

Qualifications

  • Deep knowledge of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS and MACsec.
  • 5+ years of experience in network operations and incident management.
  • Experience with AWS/Azure/GCP/OCI CSP environments.

Responsibilities

  • Engage in 24/7 global shift rotations to provide remote network support.
  • Drive operational improvements in change management and daily operations.
  • Manage and operate large-scale IP network technologies in cloud and data center environments.
  • Collaborate on network device upgrades and capacity planning.

Skills

TCP/IP
BGP
OSPF
MPLS
VxLAN
EVPN
QoS
IPsec
GRE
DNS
MACsec

Education

Bachelor's degree in Computer Science or related

Tools

Arista
Fortinet
Juniper

Job description

## Senior Network Reliability Engineer - DGX CloudApplylocations: US, CA, Santa Clara: US, Remotetime type: Full timeposted on: Posted Yesterdayjob requisition id: JR2018113NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence.In this role, the Senior Network Operations Engineer will remediate critical alerts within defined SLAs, triage production impacting network incidents, and interact with internal customers on network related issues. They will also be responsible for engaging with external vendors to remediate hardware and software issues, and participate in project related work such as network device upgrades and capacity augmentations. An ideal candidate will possess a wide range of skills, including alert monitoring & resolution in large-scale networks and CSP environments, outstanding troubleshooting skills, understanding of L3 underlay networks, and network protocol knowledge in large multi-vendor infrastructures.**What you will be doing:*** Engage in 24/7 global shift rotations to provide remote support for network repairs and changes while collaborating across teams and updating customers on status and ticket information.* Drive operational improvements in change management and daily operations by following procedures.* Manage and operate large scale IP network technologies and infrastructures.* Utilize your skills in Peering and Datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits.* Monitor and support the network health of on-premises and cloud infrastructures.* Collaborate and develop workflow enhancements while documenting best practices.**What we need to see:*** Deep knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec.* 5+ years of experience in network operations.* Skilled in network troubleshooting techniques and demonstrating creative problem-solving abilities.* Strong track record of alert response within defined SLAs and Incident management.* Experience with one or more of the following CSP environments: AWS, Azure, GCP, OCI.* Familiarity with Arista, Fortinet and Juniper.* Hands-on experience with contributing to tooling and automation for provisioning, monitoring, and managing complex network infrastructures.* Bachelor’s degree in Computer Science, related technical field, or equivalent experience.* Excellent verbal and written communication skills.**Ways To Stand Out From The Crowd:*** Solid understanding of Mellanox/Cumulus OS and Infiniband technology.* Skilled in Unix/Linux system administration, with the ability to write and understand Python/Shell scripts to improve efficiency in hyperscale environments.* Familiarity with leveraging tools such as Netbox/Nautobot, Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate about innovating and investing in ground breaking technologies.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these outstanding opportunities in deep learning cloud solutions.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 136,000 USD - 224,250 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 15, 2026.This posting is for an existing vacancy.NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Engineer - DGX Cloud
Senior Network Engineer - DGX Cloud

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Senior Software Engineer, Core Infrastructure Services - DGX Cloud

NVIDIA Corporation • United States

Remote
USD 168,000 - 322,000
Equity
Benefits
Senior Manager, GPU Cloud Infrastructure - GeForce NOW
Senior Manager, GPU Cloud Infrastructure - GeForce NOW

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 256,000 - 414,000
Solutions Architect, Ethernet Networking - NVIS
Solutions Architect, Ethernet Networking - NVIS

NVIDIA Corporation • York, Northern (KY)

Hybrid
USD 124,000 - 196,000
Equity
Benefits
Senior HPC Support Engineer - Ethernet and AI Infrastructure
Senior HPC Support Engineer - Ethernet and AI Infrastructure

NVIDIA Corporation • Memphis (TN)

Hybrid
USD 108,000 - 173,000
Comprehensive benefits package
Equity eligibility
Flexible work environment
Senior Network Engineer - DGX Cloud
Senior Network Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 265,000
Senior Customer Success Engineer - DGX Cloud
Senior Customer Success Engineer - DGX Cloud

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, Networking DGX Cloud
Senior Software Engineer, Networking DGX Cloud

NVIDIA Corporation • Town of Texas (WI)

On-site
USD 200,000 - 391,000
Equity
Benefits
Senior Full-Stack Lead Engineer
Senior Full-Stack Lead Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 258,000 - 432,000
Equity
Comprehensive benefits