Senior Software Development Engineer in Test - Datacenter Server OS

NVIDIA

Santa Clara (CA)

On-site

USD 140,000 - 270,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in Santa Clara is seeking a Platform Software QA Engineer to own and execute test plans for HGX/DGX/MGX server platforms, OS, firmware and CUDA stack.

You will build automation for server and OS testing, drive root-cause analysis, collaborate in an agile team, and contribute to reliability, CI/CD pipelines, and AI tooling validation.

Qualifications

  • Bachelor's degree in STEM or equivalent experience.
  • 5+ years experience or master's degree.
  • OS/server automation, CI/CD, and DevOps with Python, SHELL, Ansible, Jenkins, C/C++, Java, JavaScript.
  • Strong Linux troubleshooting in bare-metal and VM environments.
  • Experience with AI tools/frameworks and NLP, LLM benchmarking.
  • Experience with AI development tools for test plans and automation.
  • Experience with FW/BMC/Network/storage concepts; PCIe, IO, CPU, memory, UEFI.
  • GitHub/GitLab/Gerrit, PXE, SLURM, Kubernetes/Docker is a plus.

Responsibilities

  • Develop and execute test plans for NVIDIA HGX/DGX/MGX platform on servers, OS, firmware and CUDA stack.
  • Install and test various systems OS, server firmware and software stack.
  • Drive root-cause analysis on reliability failures and mitigation.
  • Build and debug server and OS level automation front-end and back-end framework and tests.
  • Review partner and supplier test results and prescribe additional testing as needed.
  • Work in an agile software development team with high production quality standards.
  • Manage bug lifecycle and collaborate to drive solutions.

Skills

Python
SHELL
Ansible
Jenkins
C/C++
Java
JavaScript
Linux
CI/CD
DevOps

Education

Bachelor's Degree in STEM
Master's degree

Tools

GitHub
GitLab
Gerrit
KVM
VMware
Hyper-V
Kubernetes
Docker
PXE

Job description

NVIDIA is the world leader in GPU Computing. We are passionate about markets include gaming, automotive, vision, HPC, datacenters and networking in addition to our traditional OEM business. NVIDIA is also well positioned as the ‘AI Computing Company’, and NVIDIA GPUs are the brains powering Deep Learning software frameworks, analytics, data centers, and driving autonomous vehicles. We have some of the most experienced and dedicated people in the world working for us. If you are dedicated, forward-thinking, and hard-working technical people across countries sounds exciting, this job is for you. NVIDIA is looking for an outstanding individual who thrives in a diverse work environment, has outstanding interpersonal skills and possesses a strong sense of engagement and continuous process improvement. This candidate must have enterprise server integration, strong Linux experience, reliability testing with various telemetries, scale out cluster, test plan development, track record in developing AI tools and NLP, DevOps, CI/CD experience to join our platform SWQA team.

What you’ll be doing:

  • Responsible for the development and execution of NVIDIA HGX/DGX/MGX platform test plan on servers, OS, FW and CUDA SW stack from design doc.

  • Installing and testing various systems OS, server firmware and SW stack.

  • Drive support for root cause analysis on reliability and validation test failures to identify root cause(s) and achieve mitigation.

  • Build, develop/debug server and OS level automation front-end and back-end framework and tests

  • Review partner and supplier test results and prescribe additional reliability testing on components, servers, and packaging as needed.

  • Work in an agile software development team with very high production quality standards.

  • Manage bug lifecycle and collaborate with inter-groups to drive for solutions.

What we need to see:

  • Bachelor’s Degree (or equivalent experience) in a STEM (Science, Technology, Engineering, Math or Physics) field

  • 5+ years proven experience; or master’s degree.

  • Proven years of OS and server level automation, CI/CD process and DevOps experience using Python, SHELL, Ansible, Jenkins, C/C++, Java, JavaScript

  • Strong server and Linux(Ubuntu, RedHat, CentOS, SuSE, Fedora and etc…) troubleshooting and debugging experience in a bare-metal and KVM/VMWare/Hyper-V environment.

  • Good knowledge and hands-on experience in model testing, AI tools/frameworks (TensorFlow, Pytorch, Cursor and etc…), NLP and LLM benchmarking

  • Experience in using AI development tools for test plans creation, test cases development and test cases automation

  • Strong experience in FW, BMC/OpenBMC, Network protocol, internal/external enterprise storage devices, PCIe buses and devices, IO sub-devices, CPU and memory, ACPI, UEFI spec, Redfish - huge plus

  • Proven years of experience in GitHub/Gitlab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker) – huge plus

Ways to stand out from the crowd:

  • AI related tools, LLM and NLP.

  • Experience working with NVIDIA GPU hardware is a strong plus.

  • Good to have solid understanding of virtualization in Linux (KVM, Docker orchestrated with Kubernetes)

  • Background in parallel programming ideally CUDA/OpenCL is a plus

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4.

You will also be eligible for equity and benefits.

This posting is for an existing vacancy.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Development Engineer in Test - Datacenter Server OS
Senior Software Development Engineer in Test - Datacenter Server OS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Equity
Benefits
Senior Software Development Engineer in Test - Datacenter Server OS
Senior Software Development Engineer in Test - Datacenter Server OS

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 140,000 - 224,000
Equity
Benefits
Senior Software QA Test Development Engineer - Diagnostics
Senior Software QA Test Development Engineer - Diagnostics

NVIDIA • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Equity
Benefits
Senior Software QA Test Development Engineer - Diagnostics
Senior Software QA Test Development Engineer - Diagnostics

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 140,000 - 224,250
Equity
Comprehensive benefits
Senior Software QA Test Development Engineer - Diagnostics
Senior Software QA Test Development Engineer - Diagnostics

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits
Senior Systems Performance Engineer
Senior Systems Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 136,000 - 259,000
Senior Software Engineer - Server Manageability
Senior Software Engineer - Server Manageability

2100 NVIDIA USA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits
Quality Assurance Software Developer Engineer in Test, GeForce GPU
Quality Assurance Software Developer Engineer in Test, GeForce GPU

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 140,000 - 224,250
Equity
Benefits
Senior System Software Engineer, Enterprise MODS
Senior System Software Engineer, Enterprise MODS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Systems Software Engineer, Data Center Platform Enablement
Senior Systems Software Engineer, Data Center Platform Enablement

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500