Staff Data Center Solutions Engineer - GPU Cluster Expert

Super Micro Computer Spain, S.L.

San Jose (CA)

On-site

USD 165,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Super Micro Computer is seeking a Staff Data Center Solutions Engineer to be the primary technical contact for customers deploying large-scale GPU clusters at customer sites and in partner facilities. You will guide planning, deployment, and validation with strong presentation and communication skills.

The role emphasizes advising on cooling and power products, AI/NeoCloud platforms, and internal tooling, with opportunities to lead a small team as the group scales.

Qualifications

  • Bachelor's degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Computer Science, or related field—or equivalent experience—with 12+ years of experience in datacenter infrastructure or GPU cluster deployment.
  • Strong understanding of AI/NeoCloud infrastructure architectures and GPUaaS platforms.
  • Strong understanding of how software is built; able to direct and evaluate internal tooling work without necessarily writing production code.
  • Practical experience using Agentic AI tools and workflows for diagnostics or automation.
  • Hands‑on experience planning, deploying, and validating GPU compute clusters (NVIDIA/CUDA preferred).
  • Working knowledge of BMS integration and datacenter infrastructure monitoring tools (e.g., Hyperview, Grafana, PagerDuty, or similar).
  • Solid understanding of datacenter cooling and power systems — liquid cooling (DLC), CDUs, Sidecars, cooling towers, chillers, and three‑phase power distribution.
  • Excellent presentation and communication skills; comfortable and credible presenting technical recommendations directly to customers.
  • Demonstrated ability to lead cross‑functional technical teams without formal authority (dotted-line leadership).

Responsibilities

  • Drive DataCenter Systems Planning & Deployment: Work directly with customers to plan, deploy, and validate GPU cluster deployments, coordinating infrastructure build-out, rack power-up, network fabric bring-up, cooling commissioning, and system validation.
  • Customer Advisory on Cooling & Power Products: Advise customers on Supermicro's DCBBS DLC and air-cooled product lines — CDUs, Sidecars, cooling towers, and chillers — helping them select and configure the right solution for their deployment.
  • BMS & Infrastructure Monitoring Integration: Connect and configure Building Management System (BMS) monitoring integrations alongside datacenter infrastructure monitoring tools (power, cooling, environmental) to ensure clusters are deployed into healthy, well‑instrumented environments.
  • AI/NeoCloud & GPUaaS Expertise: Apply strong working knowledge of NeoCloud infrastructure and GPUaaS platforms, including Agentic AI tooling, to guide deployment decisions, automate diagnostics, and troubleshoot issues unique to AI infrastructure.
  • Internal Tooling Direction: Partner with engineering to scope and direct internal tooling and automation using Agentic Coding tools, applying software development understanding without requiring hands‑on coding as a core function.
  • Cross-Team Technical Leadership: Diagnose complex, multi-disciplinary issues spanning power, cooling, networking, and compute, and lead (often dotted-line) teams through resolution during critical deployment windows.
  • Customer-Facing Delivery: Serve as the primary technical point of contact for customers throughout planning, deployment, and validation, managing expectations and communicating status/issues clearly.
  • Partner Datacenter Coordination: Work with partner colocation facility teams to align on power, cooling, and space readiness ahead of and during deployments.
  • AI Factory Service Strategy: Contribute to strategies and improvements for Supermicro's end-to-end AI Factory deployment services, drawing on direct customer fulfillment experience to identify gaps and opportunities.
  • Documentation & Runbooks: Develop and maintain installation runbooks, deployment checklists, and troubleshooting playbooks to support repeatable, high‑quality deployments as the team scales.

Skills

GPU cluster deployment
Data center infrastructure
AI/NeoCloud infra
Agentic AI tooling
Internal tooling direction
Cross-team leadership
BMS integration
Cooling/power systems
Presentation/communication
Customer-facing delivery

Education

Bachelor's degree in Electrical/Computer/Mechanical Engineering or CS

Tools

Hyperview
Grafana
PagerDuty
NVIDIA CUDA
InfiniBand / RoCEv2

Job description

Super Micro Computer is seeking a Staff Data Center Solutions Engineer to be the primary technical contact for customers deploying large-scale GPU clusters at customer sites and in partner facilities. You will guide planning, deployment, and validation with strong presentation and communication skills.

The role emphasizes advising on cooling and power products, AI/NeoCloud platforms, and internal tooling, with opportunities to lead a small team as the group scales.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Center GPU Solutions Engineer
Senior Data Center GPU Solutions Engineer

Supermicro • Wayne (CA)

On-site
USD 165,000 - 200,000
Senior Data Center Solutions Engineer – GPU & AI Infra
Senior Data Center Solutions Engineer – GPU & AI Infra

Support Revolution • San Jose (CA)

On-site
USD 165,000 - 200,000
Data Center Solutions Engineer: GPU & AI Infra
Data Center Solutions Engineer: GPU & AI Infra

Supermicro • San Jose (CA)

On-site
USD 165,000 - 200,000
Comprehensive benefits
Bonus and equity programs
Staff Data Center Solutions Engineer (29832)
Staff Data Center Solutions Engineer (29832)

Supermicro • San Jose (CA)

On-site
USD 165,000 - 200,000
Comprehensive benefits
Bonus and equity programs
Staff Data Center Solutions Engineer
Staff Data Center Solutions Engineer

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 165,000 - 200,000
Staff Data Center Solutions Engineer
Staff Data Center Solutions Engineer

Supermicro • Wayne (CA)

On-site
USD 165,000 - 200,000
Staff Data Center Solutions Engineer
Staff Data Center Solutions Engineer

Support Revolution • San Jose (CA)

On-site
USD 165,000 - 200,000
System Engineer - GPU Server | Data Center, Equity
System Engineer - GPU Server | Data Center, Equity

Supermicro • San Jose (CA)

On-site
USD 90,000 - 110,000
GPU Server Systems Engineer – HPC & Data Center
GPU Server Systems Engineer – HPC & Data Center

Support Revolution • San Jose (CA)

On-site
USD 90,000 - 110,000
Staff AI Data Center Solutions Lead
Staff AI Data Center Solutions Lead

Support Revolution • San Jose (CA)

On-site
USD 190,000 - 213,000