Technical Lead, Network System Validation for AI Clusters

NVIDIA

United States

A distancia

USD 180.000 - 240.000

Jornada completa

14 días+
Generador de candidaturas

Una candidatura completa en un minuto — currículum y carta de presentación adaptados, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

NVIDIA is seeking a Technical Lead to join the Network System Validation group. You will champion validation methodologies for high-speed networking in large AI cluster solutions and own the technical roadmap for a team of engineers.

You’ll read C/C++/Python code, build benchmarks and automation, and collaborate with software and hardware teams to push NCCL, RoCE, RDMA, and related components to new performance heights in production-scale AI environments.

Formación

  • B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience.
  • 8+ years of experience in networking, system validation, or related domains.
  • Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause.
  • Ability to read, debug, and reason about C/C++ code (Rust or Go a plus).
  • Strong scripting and automation experience using Python, Bash, and/or Ansible.
  • Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress.
  • Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed.
  • Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.

Responsabilidades

  • Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions
  • Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.
  • Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution
  • Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities
  • Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection
  • Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations
  • Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes

Conocimientos

Networking
System validation
C/C++ debugging
Python scripting
Distributed systems
Leadership
Test automation
Root cause analysis

Educación

B.Sc. in CS/EE

Herramientas

NCCL
RoCE
RDMA
Kubernetes

Descripción del empleo

NVIDIA is seeking a Technical Lead to join the Network System Validation group. You will champion validation methodologies for high-speed networking in large AI cluster solutions and own the technical roadmap for a team of engineers.

You’ll read C/C++/Python code, build benchmarks and automation, and collaborate with software and hardware teams to push NCCL, RoCE, RDMA, and related components to new performance heights in production-scale AI environments.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Software Engineer, Network System Validation
Senior Software Engineer, Network System Validation

NVIDIA • EE. UU.

A distancia
USD 180.000 - 240.000
Senior Networking Solutions Engineer for AI Clusters
Senior Networking Solutions Engineer for AI Clusters

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Híbrido
USD 168.000 - 322.000
Equity
Benefits package
Senior AI Cluster Network Solutions Engineer
Senior AI Cluster Network Solutions Engineer

NVIDIA • Durham (NC)

Presencial
USD 168.000 - 322.000
Equity
Benefits package
Senior Networking Solutions Engineer for AI Clusters
Senior Networking Solutions Engineer for AI Clusters

NVIDIA Corporation • Durham (CA)

Híbrido
USD 168.000 - 322.000
Senior AI Networking Performance Architect — Equity
Senior AI Networking Performance Architect — Equity

NVIDIA Corporation • Santa Clara (CA)

Presencial
USD 320.000 - 489.000
Equity
Benefits
Senior AI Cluster Network Engineer
Senior AI Cluster Network Engineer

NVIDIA • Seattle (WA)

Presencial
USD 168.000 - 322.000
Senior Networking Solutions Engineer for AI Clusters Equity
Senior Networking Solutions Engineer for AI Clusters Equity

Thomas To • Santa Clara (CA)

Presencial
USD 200.000 - 322.000
Equity
Benefits package
Senior Network Solutions Engineer (AI Clusters)
Senior Network Solutions Engineer (AI Clusters)

NVIDIA • New York (NY)

Presencial
USD 200.000 - 322.000
Equity
Comprehensive benefits
Senior Networking Solutions Engineer, AI Clusters
Senior Networking Solutions Engineer, AI Clusters

Nvidia Corporation in • Santa Clara (CA)

Presencial
USD 168.000 - 322.000
Equity
Benefits package
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

Presencial
USD 272.000 - 431.250