Senior System Engineer, HPC/AI & Cloud Clusters

Supermicro

San Jose (CA)

On-site

USD 137,000 - 156,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Supermicro is seeking a Senior System Engineer to design, deploy and validate rack-scale server solutions for data centers and enterprise customers. You will work across compute, storage, networking and HPC/AI workloads, handling testing, integration and on-site deployment.

The role emphasizes hands-on experimentation, scripting and automation, with a focus on performance benchmarks, OS/network tuning and collaboration with product teams to drive reliable, scalable infrastructure.

Qualifications

  • BS/MS in Electrical or Computer Engineering or related field; MS preferred.
  • 8+ years in server/network/storage hardware testing, debugging and troubleshooting.
  • 8+ years in DevOps or cloud environments including Docker/Containers and Kubernetes.
  • Experience with AI/ML frameworks like PyTorch or TensorFlow.
  • Familiar with TCP/IP stack, DNS/DHCP and related protocols.
  • Familiar with HPC, AI or cloud benchmark tests and networking architectures.

Responsibilities

  • Deploy Rack/Cluster infrastructure and perform comprehensive system tests on GPUs, CPUs, network and storage.
  • Lead proof-of-concept design/testing; optimize benchmarks for HPC/AI workloads.
  • Lead day-to-day support for Cluster, Storage, HPC and Cloud infra; document issues.
  • Write test procedures, reports and troubleshooting guides for servers and networks.
  • Deliver on-site deployment services to ensure customer acceptance.
  • Develop automation tools for cluster deployment and test environments.

Skills

8+ years server/hardware
8+ years DevOps/cloud
strong communication
CCNA/OpenStack/OpenShift/AWS
TCP/IP / DNS / DHCP
MLPerf / RCCL/NCCL familiarity

Education

BS/MS in Electrical/Computer Engineering or related field

Tools

Python
Shell scripting
Docker/Containers
Kubernetes
PyTorch/TensorFlow
HPC/AI benchmarks

Job description

Supermicro is seeking a Senior System Engineer to design, deploy and validate rack-scale server solutions for data centers and enterprise customers. You will work across compute, storage, networking and HPC/AI workloads, handling testing, integration and on-site deployment.

The role emphasizes hands-on experimentation, scripting and automation, with a focus on performance benchmarks, OS/network tuning and collaboration with product teams to drive reliable, scalable infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC & AI Cluster Engineer
Senior HPC & AI Cluster Engineer

Support Revolution • San Jose (CA), Northern (KY)

Hybrid
USD 137,000 - 156,000
Sr. System Engineer
Sr. System Engineer

Support Revolution • San Jose (CA), Northern (KY)

On-site
USD 137,000 - 156,000
Sr. System Engineer
Sr. System Engineer

Supermicro • San Jose (CA)

On-site
USD 137,000 - 156,000
Senior Data Center Engineer — Systems & Infrastructure
Senior Data Center Engineer — Systems & Infrastructure

Support Revolution • San Jose (CA)

On-site
USD 105,000 - 130,000
System Engineer - Data Center Benchmarks & Server Tech
System Engineer - Data Center Benchmarks & Server Tech

Support Revolution • San Jose (CA)

On-site
USD 85,000 - 100,000
Data Center Software Engineer – AI & Automation
Data Center Software Engineer – AI & Automation

Support Revolution • San Jose (CA)

On-site
USD 100,000 - 115,000
System Engineer - Data Center & Storage Benchmarks
System Engineer - Data Center & Storage Benchmarks

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 100,000 - 115,000
Staff Systems Engineer - Automation & Data Center Solutions
Staff Systems Engineer - Automation & Data Center Solutions

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 160,000 - 185,000
Staff System Engineer: CI/CD, SRE & Server/Storage
Staff System Engineer: CI/CD, SRE & Server/Storage

Supermicro • Wayne (CA)

On-site
USD 160,000 - 185,000
Bonus program
Equity
Senior GPU Platform Engineer for AI/HPC
Senior GPU Platform Engineer for AI/HPC

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 137,000 - 156,000