Senior System Engineer – GPU Platforms

Jobtailor

San Jose (CA)

On-site

USD 150,000 - 210,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in San Jose, CA seeks a Senior GPU Platform Systems Engineer to lead support, qualification, and deployment of enterprise GPU server platforms. You will diagnose complex issues, coordinate with cross-functional teams, and mentor junior engineers.

The role requires 5–15 years in systems engineering or HPC, strong Linux experience, and hands-on GPU/CUDA expertise. Preferred: NVIDIA NVQUAL and Docker/Kubernetes familiarity.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 5–15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or related field.
  • Strong knowledge of enterprise server hardware and system architecture.
  • Hands-on experience with Linux server environments.
  • Experience installing, configuring, validating, and troubleshooting server hardware and software.
  • Strong system-level troubleshooting and root-cause-analysis skills.
  • Working knowledge of PCIe architectures and high-performance I/O.
  • Experience with GPU computing, accelerators, or comparable HPC technologies.
  • Ability to independently manage complex technical assignments and drive issues toward resolution.
  • Strong written and verbal communication skills.
  • Ability to work effectively with cross-functional and geographically distributed engineering teams.
  • Comfortable participating in customer-facing technical discussions.
  • Preferred: Hands-on experience with NVIDIA data center or professional GPU platforms.
  • Preferred: Experience with CUDA and NVIDIA GPU software environments.
  • Preferred: Experience with NVIDIA NVQUAL or similar platform qualification processes.
  • Preferred: Experience with 4-GPU or 8-GPU server platforms.
  • Preferred: Familiarity with NVIDIA Blackwell, B200, Rubin or comparable accelerator architectures.
  • Preferred: Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU-to-NIC communication.
  • Preferred: Experience with GPUDirect RDMA, InfiniBand, RoCE, or high-speed Ethernet.
  • Preferred: Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools.
  • Preferred: Experience with Docker, containers, Kubernetes, or related orchestration technologies.
  • Preferred: Experience supporting AI, machine learning, HPC, or accelerated computing environments.
  • Preferred: Experience with customer POCs, technical evaluations, or engineering escalations.
  • Preferred: Experience delivering technical training or knowledge-sharing sessions.
  • Bash, Python, or other scripting experience is a plus.

Responsibilities

  • Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms.
  • Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes.
  • Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components.
  • Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation.
  • Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities.
  • Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution.
  • Collaborate with Architecture, Systems, Software, Validation, Product Management, other engineering teams, and external technology partners.
  • Develop technical documentation, troubleshooting guides, and best practices.
  • Deliver technical presentations, training sessions, and internal knowledge-sharing activities.
  • Serve as a technical resource and mentor for other engineers when appropriate.

Skills

Linux Server Environments
System-Level Troubleshooting
PCIe Architectures
High-Performance Computing
NVIDIA Data Center Platforms
Python Scripting
Bash Scripting
GPU Computing
Cross-Functional Collaboration
Customer-Facing Technical Discussions

Education

Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or related discipline

Tools

Docker
Kubernetes
NVIDIA NVQUAL
CUDA

Job description

  • Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms
  • Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes
  • Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components
  • Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation
  • Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities
  • Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution
  • Collaborate with Architecture, Systems, Software, Validation, Product Management, other engineering teams, and external technology partners
  • Develop technical documentation, troubleshooting guides, and best practices
  • Deliver technical presentations, training sessions, and internal knowledge-sharing activities
  • Serve as a technical resource and mentor for other engineers when appropriate
Requirements
  • Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or a related discipline, or equivalent practical experience
  • 5–15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or a related field
  • Strong knowledge of enterprise server hardware and system architecture
  • Hands-on experience with Linux server environments
  • Experience installing, configuring, validating, and troubleshooting server hardware and software
  • Strong system-level troubleshooting and root-cause-analysis skills
  • Working knowledge of PCIe architectures and high-performance I/O
  • Experience with GPU computing, accelerators, or comparable high-performance computing technologies
  • Ability to independently manage complex technical assignments and drive issues toward resolution
  • Strong written and verbal communication skills
  • Ability to work effectively with cross-functional and geographically distributed engineering teams
  • Comfortable participating in customer-facing technical discussions
  • Preferred: Hands-on experience with NVIDIA data center or professional GPU platforms
  • Preferred: Experience with CUDA and NVIDIA GPU software environments
  • Preferred: Experience with NVIDIA NVQUAL or similar platform qualification processes
  • Preferred: Experience with 4-GPU or 8-GPU server platforms
  • Preferred: Familiarity with NVIDIA Blackwell, B200, Rubin, or comparable accelerator architectures
  • Preferred: Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU-to-NIC communication
  • Preferred: Experience with GPUDirect RDMA, InfiniBand, RoCE, or high-speed Ethernet
  • Preferred: Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools
  • Preferred: Experience with Docker, containers, Kubernetes, or related orchestration technologies
  • Preferred: Experience supporting AI, machine learning, HPC, or accelerated computing environments
  • Preferred: Experience with customer POCs, technical evaluations, or engineering escalations
  • Preferred: Experience delivering technical training or knowledge-sharing sessions
  • Bash, Python, or other scripting experience is a plus
Core Competencies

Demonstrates expertise in GPU server platform support, including installation, configuration, and troubleshooting of Linux environments and GPU technologies. Proficient in system-level diagnostics, technical documentation, and cross-functional collaboration to drive complex technical assignments to resolution.

Highest-signal resume keywords
  • GPU Computing
  • Linux Server Environments
  • System-Level Troubleshooting
  • NVIDIA NVQUAL
  • Technical Documentation
Hard Skills
  • GPU Drivers
  • CUDA
  • Server Hardware Configuration
  • Root-Cause Analysis
  • PCIe Architectures
  • High-Performance Computing
  • NVIDIA Data Center Platforms
  • Docker
  • Python Scripting
  • Bash Scripting
Soft Skills
  • Strong Communication Skills
  • Cross-Functional Collaboration
  • Customer-Facing Technical Discussions
  • Mentoring
Industry Keywords
  • Systems Engineering
  • Server Engineering
  • Platform Engineering
  • Technical Enablement
  • AI Infrastructure
  • HPC
Tools & Technologies
  • NCCL
  • NVML
  • DCGM
  • Fabric Manager
  • InfiniBand
  • RoCE
  • High-Speed Ethernet
  • Kubernetes
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer 4
GPU Systems Engineer 4

RPMGlobal • Bethesda (MD)

On-site
USD 140,000 - 190,000
Sr. System Engineer/GPU Platforms
Sr. System Engineer/GPU Platforms

Supermicro • Wayne (CA)

On-site
USD 137,000 - 156,000
Senior Systems Software Engineer, Data Center Platform Enablement
Senior Systems Software Engineer, Data Center Platform Enablement

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity options
Comprehensive benefits
Inclusive work environment
Sr. System Engineer/GPU Platforms
Sr. System Engineer/GPU Platforms

Support Revolution • San Jose (CA)

On-site
USD 137,000 - 156,000
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior Linux Kernel Systems Software Engineer – CSP Engagements
Senior Linux Kernel Systems Software Engineer – CSP Engagements

NVIDIA • Redmond (WA)

On-site
USD 130,000 - 160,000
Senior System Software Engineer - GPU Server
Senior System Software Engineer - GPU Server

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Generous benefits package
Equity eligibility
GPU Systems Engineer 3
GPU Systems Engineer 3

RPMGlobal • Bethesda (MD)

On-site
USD 120,000 - 260,000
Senior Datacenter Product Development Engineer
Senior Datacenter Product Development Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 259,000
Equity options
Competitive benefits package
Senior System Software Engineer - GPU Server
Senior System Software Engineer - GPU Server

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000