Systems Software Engineer, Kubernetes Scale - DGX Cloud

NVIDIA

France

Sur place

EUR 60 000 - 90 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

NVIDIA is looking for an innovative engineer to drive performance and scale characterization for the NVIDIA DGX Cloud software stack. The role involves collaborating with researchers and developers, performing automated testing, and solving complex issues related to Kubernetes.

The ideal candidate has a Bachelor's or Master's degree in Engineering, 2+ years of experience, and is proficient in Kubernetes and programming languages like Golang or Python.

This position offers a unique opportunity to work with cutting-edge technology and contribute to impactful AI solutions.

Qualifications

  • 2+ years of experience in Computer Architecture, Networking, or Storage systems.
  • Expertise in Kubernetes and familiarity with related CNCF projects.
  • Experience with performance modeling and benchmarking at scale.

Responsabilités

  • Drive performance and scale characterization for the NVIDIA DGX Cloud software stack.
  • Collaborate with researchers and developers to develop automated tests.
  • Debug issues related to operating Kubernetes clusters at ultra-large scale.

Connaissances

Kubernetes
Golang/Python
Performance modeling
Distributed systems
AI workload optimization

Formation

Bachelor's/Master's in Engineering

Outils

NVIDIA software ecosystem

Description du poste

The DGX Cloud organization at NVIDIA brings together cutting‑edge hardware and software innovation to deliver industry‑leading accelerated computing for the world’s most adventurous AI workloads. We’re a team of innovative engineers dedicated to solving some of the world’s biggest challenges, constantly driving advancements and impacting millions of lives worldwide.

What You’ll Be Doing
  • Drive end‑to‑end performance and scale characterization for the NVIDIA DGX Cloud software stack, from Kubernetes control and data planes through NVIDIA components such as GPU Operator, Network Operator, DCGM, NIM, and distributed inference serving, following issues from orchestration down to the metal.
  • Collaborate with AI researchers, developers and customers to develop innovative, automated tests that simulate real user workloads using custom‑built and leading open‑source tools and frameworks.
  • Deep dive into performance and scale issues in complex distributed systems, including interactions between Kubernetes and the NVIDIA software stack, to identify and resolve root causes.
  • Design and develop monitoring, reporting and analysis tools for performance and scale testing across software, GPU and CPU resources.
  • Triage, debug and root cause issues related to operating Kubernetes clusters at ultra‑large scale, ensuring reliability and efficiency.
  • Build and maintain a high‑velocity framework that enables continuous, always‑on performance and scale testing via a modern CI/CD pipeline.
  • Document research, methodologies and results clearly and concisely, and present findings at internal and external venues, including community conferences such as KubeCon and GTC.
  • Engage efficiently with upstream communities — including Kubernetes, CNCF and NVIDIA open‑source projects — to validate performance and scalability of AI workloads early and help shape design and development decisions.
What We Need To See
  • 2+ years of experience in Computer Architecture, Networking, Storage systems, Accelerators and Bachelors/Masters in Engineering (preferably Electrical Engineering, Computer Engineering, or Computer Science) or equivalent experience.
  • Expertise in Kubernetes and familiarity with related CNCF projects.
  • Background in working with large‑scale parallel and distributed accelerator‑based systems.
  • Expertise optimizing performance and AI workloads on large‑scale systems.
  • Experience with performance modeling and benchmarking at scale.
  • Proficiency in Golang/Python.
  • Background with the NVIDIA software ecosystem in both training and inference domains.
  • Expertise with at least one of public CSP infrastructure (GCP, AWS, Azure, OCI, for example).
Ways To Stand Out From The Crowd
  • Strong operational experience with any one of the Kubernetes distributions.
  • Prior experience scaling Kubernetes clusters to ultra‑large node and object counts.
  • Demonstrated history of working in the open‑source community.
  • Excellent communication and interpersonal abilities.
  • PhD in relevant areas.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people in the world working for us. If you’re creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 176,250 PLN – 305,500 PLN for Level 2, and 221,250 PLN – 383,500 PLN for Level 3. JR2020236

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior AI Compute Engineer
Senior AI Compute Engineer

NVIDIA • France

Sur place
EUR 90 000 - 150 000
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • France

Sur place
EUR 90 000 - 150 000
Senior AIC Engineer
Senior AIC Engineer

NVIDIA • Marseille

Sur place
EUR 90 000 - 120 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA AI • Aillas

Sur place
EUR 120 000 - 180 000
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA Corporation • France

À distance
EUR 90 000 - 120 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • France

Sur place
EUR 110 000 - 160 000
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA • Courbevoie

Sur place
EUR 90 000 - 140 000
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure
Solutions Architect - NVIDIA AI Cloud Partners and Datacentre Infrastructure

NVIDIA Gruppe • Courbevoie

Sur place
EUR 120 000 - 180 000
Data Center Engineer
Data Center Engineer

Trust In SODA • Paris

Sur place
EUR 60 000 - 90 000
General Account Software Sales Specialist, General Account Software Sales Specialist
General Account Software Sales Specialist, General Account Software Sales Specialist

NVIDIA • Courbevoie

Sur place
EUR 55 000 - 80 000
Comprehensive benefits package
Competitive salaries