Senior GPU Platform Engineer – Linux, NVQUAL & HPC

Jobtailor

San Jose (CA)

On-site

USD 150,000 - 210,000

Full time

13 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Jobtailor in San Jose, CA seeks a Senior GPU Platform Systems Engineer to lead support, qualification, and deployment of enterprise GPU server platforms. You will diagnose complex issues, coordinate with cross-functional teams, and mentor junior engineers.

The role requires 5–15 years in systems engineering or HPC, strong Linux experience, and hands-on GPU/CUDA expertise. Preferred: NVIDIA NVQUAL and Docker/Kubernetes familiarity.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 5–15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or related field.
  • Strong knowledge of enterprise server hardware and system architecture.
  • Hands-on experience with Linux server environments.
  • Experience installing, configuring, validating, and troubleshooting server hardware and software.
  • Strong system-level troubleshooting and root-cause-analysis skills.
  • Working knowledge of PCIe architectures and high-performance I/O.
  • Experience with GPU computing, accelerators, or comparable HPC technologies.
  • Ability to independently manage complex technical assignments and drive issues toward resolution.
  • Strong written and verbal communication skills.
  • Ability to work effectively with cross-functional and geographically distributed engineering teams.
  • Comfortable participating in customer-facing technical discussions.
  • Preferred: Hands-on experience with NVIDIA data center or professional GPU platforms.
  • Preferred: Experience with CUDA and NVIDIA GPU software environments.
  • Preferred: Experience with NVIDIA NVQUAL or similar platform qualification processes.
  • Preferred: Experience with 4-GPU or 8-GPU server platforms.
  • Preferred: Familiarity with NVIDIA Blackwell, B200, Rubin or comparable accelerator architectures.
  • Preferred: Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU-to-NIC communication.
  • Preferred: Experience with GPUDirect RDMA, InfiniBand, RoCE, or high-speed Ethernet.
  • Preferred: Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools.
  • Preferred: Experience with Docker, containers, Kubernetes, or related orchestration technologies.
  • Preferred: Experience supporting AI, machine learning, HPC, or accelerated computing environments.
  • Preferred: Experience with customer POCs, technical evaluations, or engineering escalations.
  • Preferred: Experience delivering technical training or knowledge-sharing sessions.
  • Bash, Python, or other scripting experience is a plus.

Responsibilities

  • Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms.
  • Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes.
  • Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components.
  • Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation.
  • Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities.
  • Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution.
  • Collaborate with Architecture, Systems, Software, Validation, Product Management, other engineering teams, and external technology partners.
  • Develop technical documentation, troubleshooting guides, and best practices.
  • Deliver technical presentations, training sessions, and internal knowledge-sharing activities.
  • Serve as a technical resource and mentor for other engineers when appropriate.

Skills

Linux Server Environments
System-Level Troubleshooting
PCIe Architectures
High-Performance Computing
NVIDIA Data Center Platforms
Python Scripting
Bash Scripting
GPU Computing
Cross-Functional Collaboration
Customer-Facing Technical Discussions

Education

Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or related discipline

Tools

Docker
Kubernetes
NVIDIA NVQUAL
CUDA

Job description

Jobtailor in San Jose, CA seeks a Senior GPU Platform Systems Engineer to lead support, qualification, and deployment of enterprise GPU server platforms. You will diagnose complex issues, coordinate with cross-functional teams, and mentor junior engineers.

The role requires 5–15 years in systems engineering or HPC, strong Linux experience, and hands-on GPU/CUDA expertise. Preferred: NVIDIA NVQUAL and Docker/Kubernetes familiarity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Platform Engineer for AI/HPC
Senior GPU Platform Engineer for AI/HPC

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 137,000 - 156,000
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits
Senior GPU Platform Engineer
Senior GPU Platform Engineer

Supermicro • Wayne (CA)

On-site
USD 137,000 - 156,000
Sr. System Engineer/GPU Platforms
Sr. System Engineer/GPU Platforms

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 137,000 - 156,000
Senior System Engineer – GPU Platforms
Senior System Engineer – GPU Platforms

Jobtailor • San Jose (CA)

On-site
USD 150,000 - 210,000
GPU Systems Infrastructure Engineer
GPU Systems Infrastructure Engineer

Blue Signal Search • Fremont (CA)

On-site
USD 120,000 - 170,000
Senior GPU HPC Systems Engineer
Senior GPU HPC Systems Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits
Senior System Software Engineer – GPU & HPC, Equity
Senior System Software Engineer – GPU & HPC, Equity

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior GPU Cloud Platform Support Engineer
Senior GPU Cloud Platform Support Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 108,000 - 173,000
Senior Hybrid Quantum–HPC Platform Engineer
Senior Hybrid Quantum–HPC Platform Engineer

NVIDIA • Connecticut

Hybrid
USD 224,000 - 431,000
Equity
Benefits package