Senior Software Engineer, Network System Validation

NVIDIA Corporation

Grézet-Cavagnan

Hybrid

EUR 90,000 - 150,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA is seeking a Technical Lead for the Network System Validation group to guide validation directions for high-speed networking across AI cluster environments. You will mentor engineers, own the validation roadmap, and push networking throughput to scale.

This hands-on leadership role combines validation methodology development with automation, debugging, and performance analysis across NCCL, RoCE, RDMA, and software/hardware stacks. Collaboration with cross-functional teams is essential.

Qualifications

  • B.Sc./B.A. in Computer Science, Electrical Engineering, or equivalent experience.
  • 8+ years of experience in networking, system validation, or related domains.
  • Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.
  • Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).
  • Strong scripting and automation experience using Python, Bash, and/or Ansible.
  • Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress.
  • Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.
  • Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.

Responsibilities

  • Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions.
  • Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.
  • Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution.
  • Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities.
  • Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection.
  • Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations.
  • Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes.

Skills

Networking
Debugging complex systems
Python scripting
Bash scripting
Ansible
Code reading (C/C++/Python)
Distributed systems
Technical leadership

Education

B.Sc./B.A. in CS or EE

Tools

NCCL
RoCE
RDMA
Logging/instrumentation
Profiling

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect — Ethernet, InfiniBand, NVLink, and BlueField DPUs — that switches thousands of GPUs into a single AI supercomputer, moving data at the scale and speed the most demanding workloads require. We are looking for a Technical Lead to join our Network System Validation group. You will work on validating advanced networking solutions across NVIDIA complex AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove our network's correctness and performance at scale. In this role you will lead the validation direction and engineering excellence of one of our technology validation teams. This is a deeply hands‑on role for a technology leader who can own the technical roadmap, technology mentoring a team of high-performance engineers, and push NVIDIA's network to its speed‑of‑light limits. The role combines the development of validation methodologies and automation tools with hands‑on debugging, performance analysis, and investigation of cutting‑edge AI networking technologies at scale.

What you’ll be doing:
  • Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions
  • Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis
  • Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution
  • Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities
  • Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection
  • Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations
  • Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes
What we need to see:
  • B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience
  • 8+ years of experience in networking, system validation, or related domains
  • Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause
  • Ability to read, debug, and reason about C/C++ code (Rust or Go a plus)
  • Strong scripting and automation experience using Python, Bash, and/or Ansible
  • Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress
  • Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed
  • Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines
Ways to stand out from the crowd:
  • Experience with large-scale clusters or distributed systems
  • Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField)
  • Background in performance analysis, Kubernetes, or cloud environments
  • Background in chaos testing, fault injection, or simulation systems

We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you! NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, Network Systems Validation for AI Clusters
Tech Lead, Network Systems Validation for AI Clusters

NVIDIA Corporation • Grézet-Cavagnan

Hybrid
EUR 90,000 - 150,000
Software Verification Engineer
Software Verification Engineer

NVIDIA Corporation • Grézet-Cavagnan

Hybrid
EUR 55,000 - 85,000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • Courbevoie

On-site
EUR 120,000 - 170,000
Solutions Architect - NVIDIA AI Cloud Partners
Solutions Architect - NVIDIA AI Cloud Partners

NVIDIA Corporation • Courbevoie

On-site
EUR 90,000 - 120,000
Senior Performance Engineer
Senior Performance Engineer

NVIDIA Corporation • Grézet-Cavagnan

Hybrid
EUR 90,000 - 120,000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA Corporation • France

Hybrid
EUR 120,000 - 180,000
Solutions Architect - NVIDIA AI Cloud Partners
Solutions Architect - NVIDIA AI Cloud Partners

NVIDIA • Courbevoie

On-site
EUR 90,000 - 130,000
Automated Verification Engineer - Networking & Cloud Infra
Automated Verification Engineer - Networking & Cloud Infra

NVIDIA Corporation • Grézet-Cavagnan

Hybrid
EUR 55,000 - 85,000
Solutions Architect - NVIDIA AI Cloud Partners
Solutions Architect - NVIDIA AI Cloud Partners

NVIDIA Corporation • France

Hybrid
EUR 90,000 - 130,000
Tech Engagement Lead, AI Labs - EMEA
Tech Engagement Lead, AI Labs - EMEA

NVIDIA • France

On-site
EUR 120,000 - 160,000