Senior Software Development Engineer in Test - Datacenter Server OS

NVIDIA AI

Santa Clara (CA)

On-site

USD 180,000 - 255,000

Full time

31 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA AI in Santa Clara, CA is seeking an experienced software quality/automation engineer to develop and execute platform test plans for HGX/DGX/MGX servers, OS, firmware and CUDA stacks. You will implement automation for server and OS testing, perform root-cause analysis, and drive reliability improvements in an agile, high-quality environment.

The role requires strong Linux troubleshooting, CI/CD/DevOps experience, and hands-on work with AI tools, NLP benchmarks, and modern virtualization

Qualifications

  • Bachelor’s Degree (or equivalent) in a STEM field.
  • 5+ years proven experience; or master’s degree.
  • OS and server level automation, CI/CD, and DevOps using Python, SHELL, Ansible, Jenkins, C/C++, Java, JavaScript.
  • Strong Linux troubleshooting in bare-metal and VM environments.
  • Experience with AI tools/frameworks and NLP/LLM benchmarking.
  • GitHub/GitLab/Gerrit, PXE, SLURM, Kubernetes/Docker experience is a plus.

Responsibilities

  • Develop and execute platform test plans for NVIDIA HGX/DGX/MGX on servers, OS, firmware and CUDA.
  • Install and test various systems OS, server firmware and SW stack.
  • Drive root-cause analysis for reliability and validation failures and implement mitigations.
  • Build/test automation front-end and back-end frameworks and tests.
  • Review tests and prescribe additional reliability testing as needed.
  • Work in an agile team with high production quality standards.

Skills

OS automation
CI/CD
Python
C/C++
Linux troubleshooting
Telemetries
CI/CD tooling
GitHub/GitLab
Docker/Kubernetes
Test automation

Education

Bachelor’s Degree in STEM
Master’s degree (advantage)

Tools

GitHub
GitLab
Gerrit
PXE
SLURM
Kubernetes
Docker

Job description

Job Requisition ID JR2011122

Job Category Engineering

Time Type Full time

NVIDIA is the world leader in GPU Computing. We are passionate about markets include gaming, automotive, vision, HPC, datacenters and networking in addition to our traditional OEM business. NVIDIA is also well positioned as the ‘AI Computing Company’, and NVIDIA GPUs are the brains powering Deep Learning software frameworks, analytics, data centers, and driving autonomous vehicles. We have some of the most experienced and dedicated people in the world working for us. If you are dedicated, forward-thinking, and hard-working technical people across countries sounds exciting, this job is for you. NVIDIA is looking for an outstanding individual who thrives in a diverse work environment, has outstanding interpersonal skills and possesses a strong sense of engagement and continuous process improvement. This candidate must have enterprise server integration, strong Linux experience, reliability testing with various telemetries, scale out cluster, test plan development, track record in developing AI tools and NLP, DevOps, CI/CD experience to join our platform SWQA team.

What You’ll Be Doing
  • Responsible for the development and execution of NVIDIA HGX/DGX/MGX platform test plan on servers, OS, FW and CUDA SW stack from design doc.
  • Installing and testing various systems OS, server firmware and SW stack.
  • Drive support for root cause analysis on reliability and validation test failures to identify root cause(s) and achieve mitigation.
  • Build, develop/debug server and OS level automation front-end and back-end framework and tests
  • Review partner and supplier test results and prescribe additional reliability testing on components, servers, and packaging as needed.
  • Work in an agile software development team with very high production quality standards.
  • Manage bug lifecycle and collaborate with inter-groups to drive for solutions.
What We Need To See
  • Bachelor’s Degree (or equivalent experience) in a STEM (Science, Technology, Engineering, Math or Physics) field
  • 5+ years proven experience; or master’s degree.
  • Proven years of OS and server level automation, CI/CD process and DevOps experience using Python, SHELL, Ansible, Jenkins, C/C++, Java, JavaScript
  • Strong server and Linux(Ubuntu, RedHat, CentOS, SuSE, Fedora and etc…) troubleshooting and debugging experience in a bare-metal and KVM/VMWare/Hyper-V environment.
  • Good knowledge and hands-on experience in model testing, AI tools/frameworks (TensorFlow, Pytorch, Cursor and etc…), NLP and LLM benchmarking
  • Experience in using AI development tools for test plans creation, test cases development and test cases automation
  • Strong experience in FW, BMC/OpenBMC, Network protocol, internal/external enterprise storage devices, PCIe buses and devices, IO sub-devices, CPU and memory, ACPI, UEFI spec, Redfish - huge plus
  • Proven years of experience in GitHub/Gitlab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker) – huge plus
Ways To Stand Out From The Crowd
  • AI related tools, LLM and NLP.
  • Experience working with NVIDIA GPU hardware is a strong plus.
  • Good to have solid understanding of virtualization in Linux (KVM, Docker orchestrated with Kubernetes)
  • Background in parallel programming ideally CUDA/OpenCL is a plus

#SDETSChiring

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 13, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Development Engineer in Test - Datacenter Server OS
Senior Software Development Engineer in Test - Datacenter Server OS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Equity
Benefits
Senior Software Development Engineer in Test - Datacenter Server OS
Senior Software Development Engineer in Test - Datacenter Server OS

NVIDIA • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Equity
Benefits
Senior Software QA Test Development Engineer - Diagnostics
Senior Software QA Test Development Engineer - Diagnostics

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 140,000 - 270,000
Equity
Benefits
Senior Software QA Test Development Engineer - Diagnostics
Senior Software QA Test Development Engineer - Diagnostics

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 140,000 - 224,250
Equity
Comprehensive benefits
Senior Systems Performance Engineer
Senior Systems Performance Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 136,000 - 259,000
Senior Systems Performance Engineer
Senior Systems Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 136,000 - 259,000
Senior Automation and Tools Development Engineer
Senior Automation and Tools Development Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Equity
Benefits
Senior Hardware Validation Engineer
Senior Hardware Validation Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 136,000 - 259,000
Senior System Test Engineer
Senior System Test Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 160,000 - 270,000
Senior System Test Engineer
Senior System Test Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 152,000 - 288,000