Sr. System Engineer/GPU Platforms

Support Revolution

San Jose (CA)

On-site

USD 137,000 - 156,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Supermicro is seeking an experienced Senior Systems Engineer / GPU Platforms to support bring‑up, qualification, and customer deployment of multi‑GPU server platforms for AI, HPC, and enterprise workloads.

The candidate will work across the product lifecycle, partner with Architecture, Systems, Software, and Validation teams, and provide technical leadership, documentation, and training to peers and customers.

Qualifications

  • Bachelor's degree in Engineering/CS/IT or equivalent practical experience.
  • 5–15 years of relevant industry experience in systems engineering, server engineering, validation or AI infrastructure.
  • Strong knowledge of enterprise server hardware and system architecture.
  • Hands‑on experience with Linux server environments.
  • Experience installing, configuring, validating, and troubleshooting server hardware and software.
  • Strong system‑level troubleshooting and root‑cause analysis skills.
  • Working knowledge of PCIe architectures and high‑performance I/O.
  • Experience with GPU computing, accelerators, or HPC technologies.
  • Ability to manage complex technical assignments and drive issues to resolution.
  • Strong written and verbal communication skills.
  • Ability to work with cross‑functional and distributed teams.
  • Comfortable participating in customer‑facing technical discussions.

Responsibilities

  • Assist bring‑up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms.
  • Execute GPU platform qualification activities including NVQUAL or equivalent.
  • Install, configure, troubleshoot Linux, GPU drivers, CUDA, firmware, libraries and related software.
  • Diagnose complex system issues using logs, telemetry, diagnostics and vendor tools.
  • Support multi‑GPU server platforms through qualification, launch and post‑release activities.
  • Engage in customer‑facing POC/EVAL engagements with system preparation and debugging.
  • Collaborate with Architecture, Systems, Software, Validation, Product Management and partners.
  • Develop technical documentation, troubleshooting guides and best practices.
  • Deliver technical presentations and training sessions; mentor others as needed.
  • Maintain ownership and accountability for complex technical issues.
  • Learn new server/GPU/software technologies quickly and effectively.
  • Support cross‑functional collaboration and knowledge sharing.

Skills

Linux server environments
Server hardware knowledge
System‑level troubleshooting
Cross‑functional collaboration
Written and verbal communication

Education

Bachelor's degree in Engineering/CS/IT

Tools

NVIDIA GPU software
CUDA
NVQUAL
Docker
Kubernetes
PCIe architectures knowledge

Job description

About Supermicro:

Supermicro® is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.

Job Summary:

Supermicro is seeking an experienced Senior Systems Engineer / GPU Platforms to support the bring-up, qualification, enablement, and customer deployment of advanced GPU computing platforms.

This role focuses on multi-GPU server systems used for AI, HPC, enterprise computing, and accelerated workloads. The successful candidate will work across the product lifecycle, from initial system bring-up and qualification through product release, customer POC/EVAL support, debugging, and post-launch technical enablement.

The ideal candidate combines strong server hardware knowledge with hands-on Linux and GPU software experience and can independently troubleshoot complex issues across hardware, firmware, operating systems, networking, and GPU software environments.

Key Responsibilities
  • Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms.
  • Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes.
  • Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components.
  • Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation.
  • Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities.
  • Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution.
  • Collaborate with internal Architecture, Systems, Software, Validation, Product Management, and other engineering teams, as well as external technology partners.
  • Develop technical documentation, troubleshooting guides, and best practices.
  • Deliver technical presentations, training sessions, and internal knowledge-sharing activities.
  • Serve as a technical resource and mentor for other engineers when appropriate.
  • Strong technical depth and systems-level troubleshooting ability.
  • Ownership and accountability for complex technical issues.
  • Ability to quickly learn new server, GPU, and software technologies.
  • Effective cross-functional collaboration.
  • Clear technical communication and documentation.
  • Willingness to share knowledge and support team development

5–15 years of relevant experience

Candidates should demonstrate the ability to independently support complex GPU platforms and technical customer environments. More senior candidates should additionally bring broad system-level expertise, technical leadership, mentoring experience, and ownership of complex platform or customer-facing initiatives.

Required Qualifications
  • Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or a related discipline, or equivalent practical experience.
  • 5–15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or a related field.
  • Strong knowledge of enterprise server hardware and system architecture.
  • Hands-on experience with Linux server environments.
  • Experience installing, configuring, validating, and troubleshooting server hardware and software.
  • Strong system-level troubleshooting and root‑cause‑analysis skills.
  • Working knowledge of PCIe architectures and high‑performance I/O.
  • Experience with GPU computing, accelerators, or comparable high‑performance computing technologies.
  • Ability to independently manage complex technical assignments and drive issues toward resolution.
  • Strong written and verbal communication skills.
  • Ability to work effectively with cross‑functional and geographically distributed engineering teams.
  • Comfortable participating in customer‑facing technical discussions.
Preferred Qualifications
  • Hands‑on experience with NVIDIA data center or professional GPU platforms.
  • Experience with CUDA and NVIDIA GPU software environments.
  • Experience with NVIDIA NVQUAL or similar platform qualification processes.Experience with 4-GPU or 8-GPU server platforms.
  • Familiarity with NVIDIA Blackwell, B200, Rubin, or comparable accelerator architectures.
  • Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU‑to‑NIC communication.
  • Experience with GPUDirect RDMA, InfiniBand, RoCE, or high‑speed Ethernet.
  • Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools.
  • Experience with Docker, containers, Kubernetes, or related orchestration technologies.
  • Experience supporting AI, machine learning, HPC, or accelerated computing environments.
  • Experience with customer POCs, technical evaluations, or engineering escalations.
  • Experience delivering technical training or knowledge‑sharing sessions.
  • Bash, Python, or other scripting experience is a plus.
Salary Range

$137,000 - $156,000

The salary offered will depend on several factors, including your location, level, education, training, specific skills, years of experience, and comparison to other employees already in this role. In addition to a comprehensive benefits package, candidates may be eligible for other forms of compensation, such as participation in bonus and equity award programs.

EEO Statement

Supermicro is an Equal Opportunity Employer and embraces diversity in our employee population. It is the policy of Supermicro to provide equal opportunity to all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status or special disabled veteran, marital status, pregnancy, genetic information, or any other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. System Engineer/GPU Platforms
Sr. System Engineer/GPU Platforms

Supermicro • Wayne (CA)

On-site
USD 137,000 - 156,000
System Engineer, GPU Server
System Engineer, GPU Server

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 90,000 - 110,000
System Engineer, GPU Server
System Engineer, GPU Server

Supermicro • San Jose (CA)

On-site
USD 90,000 - 110,000
Senior GPU Platform Engineer
Senior GPU Platform Engineer

Supermicro • Wayne (CA)

On-site
USD 137,000 - 156,000
Staff Data Center Solutions Engineer (29832)
Staff Data Center Solutions Engineer (29832)

Supermicro • San Jose (CA)

On-site
USD 165,000 - 200,000
Comprehensive benefits
Bonus and equity programs
Service Engineer
Service Engineer

Supermicro • Miami (FL)

On-site
USD 79,000 - 87,000
Senior GPU Systems Engineer: AI & HPC Platforms
Senior GPU Systems Engineer: AI & HPC Platforms

Support Revolution • San Jose (CA)

On-site
USD 137,000 - 156,000
Service Engineer
Service Engineer

Support Revolution • Columbus (OH)

On-site
USD 79,000 - 87,000
Staff Data Center Solutions Engineer
Staff Data Center Solutions Engineer

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 165,000 - 200,000
Sr. Product Manager (28901)
Sr. Product Manager (28901)

Supermicro • San Jose (CA)

On-site
USD 155,000 - 185,000