Staff Data Center Solutions Engineer

Supermicro

Wayne (CA)

On-site

USD 165,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Supermicro seeks a Staff Data Center Solutions Engineer to be the primary technical contact for planning, deploying, and validating large-scale GPU clusters at customer sites and partner facilities.

You will advise on DLC and air-cooled product lines, coordinate BMS integrations, and apply AI tooling knowledge to improve configurations and diagnostics. Strong presentation skills are essential.

Qualifications

  • Bachelor's degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Computer Science, or related field; 12+ years in datacenter infrastructure or GPU cluster deployment.
  • Strong understanding of AI/NeoCloud infrastructure architectures and GPUaaS platforms.
  • Strong understanding of how software is built; able to direct and evaluate internal tooling work without necessarily writing production code.
  • Practical experience using Agentic AI tools and workflows for diagnostics or automation.
  • Hands-on experience planning, deploying, and validating GPU compute clusters (NVIDIA/CUDA preferred).
  • Working knowledge of BMS integration and datacenter infrastructure monitoring tools (e.g., Hyperview, Grafana, PagerDuty, or similar).
  • Solid understanding of datacenter cooling and power systems—liquid cooling (DLC), CDUs, Sidecars, cooling towers, chillers, and three-phase power distribution.
  • Excellent presentation and communication skills; credible presenting technical recommendations to customers.
  • Demonstrated ability to lead cross-functional technical teams without formal authority (dotted-line).

Responsibilities

  • Drive DataCenter Systems Planning & Deployment with customers for GPU clusters.
  • Advise customers on Supermicro's DLC and air-cooled product lines (CDUs, Sidecars, cooling towers, chillers).
  • Integrate BMS and infrastructure monitoring to ensure healthy deployments.
  • Apply knowledge of NeoCloud and GPUaaS to guide deployment decisions and automate diagnostics.
  • Guide internal tooling and automation using Agentic Coding tools; minimal coding required.
  • Lead cross-disciplinary teams to resolve issues during critical deployment windows.
  • Serve as primary technical contact for customers throughout planning, deployment, and validation.
  • Coordinate with partner datacenters on power, cooling, and space readiness.
  • Contribute to AI Factory service strategies and improvements based on customer experience.
  • Create and maintain runbooks, checklists, and troubleshooting playbooks for repeatable deployments.

Skills

AI/NeoCloud infra
Leadership & cross‑functional
GPU clusters deployment
BMS integration knowledge
Presentation skills

Education

Bachelor's degree in Electrical/Computer Engineering or related field

Tools

Agentic AI tooling
Monitoring tools (Grafana, etc.)

Job description

About Supermicro:

Supermicro is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing company among the Silicon Valley Top 50 technology firms. Our unprecedented global expansion has provided us with the opportunity to offer a large number of new positions to the technology community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.

Job Summary:

As Staff Data Center Solutions Engineer you will be the primary technical point of contact from the Data Center side for customers planning, deploying, and validating large-scale GPU clusters - at customer sites, in Supermicro-owned facilities, and in partner colocation environments. This is a fulfillment role serving customers directly, day one as an individual contributor, with a clear path to leading a small team as the group scales.

This role suits a Staff Data Center Solutions Engineer who wants to grow into broad GPUaaS solution engineering while advising customers on Supermicro's DCBBS liquid-cooled (DLC) and air-cooled product lines - CDUs, Sidecars, cooling towers, and chillers. You don't need to be a professional software developer, but you should understand how software is built well enough to guide internal tooling work and speak credibly to technical customers and internal engineering teams alike. Strong presentation and communication skills are essential - you'll regularly be the person customers trust in the room.

Essential Duties and Responsibilities:

Drive DataCenter Systems Planning & Deployment: Work directly with customers to plan, deploy, and validate GPU cluster deployments, coordinating infrastructure build-out, rack power-up, network fabric bring-up, cooling commissioning, and system validation.

Customer Advisory on Cooling & Power Products: Advise customers on Supermicro's DCBBS DLC and air-cooled product lines - CDUs, Sidecars, cooling towers, and chillers - helping them select and configure the right solution for their deployment.

BMS & Infrastructure Monitoring Integration: Connect and configure Building Management System (BMS) monitoring integrations alongside datacenter infrastructure monitoring tools (power, cooling, environmental) to ensure clusters are deployed into healthy, well-instrumented environments.

AI/NeoCloud & GPUaaS Expertise: Apply strong working knowledge of NeoCloud infrastructure and GPUaaS platforms, including Agentic AI tooling, to guide deployment decisions, automate diagnostics, and troubleshoot issues unique to AI infrastructure.

Internal Tooling Direction: Partner with engineering to scope and direct internal tooling and automation using Agentic Coding tools, applying software development understanding without requiring hands-on coding as a core function.

Cross-Team Technical Leadership: Diagnose complex, multi-disciplinary issues spanning power, cooling, networking, and compute, and lead (often dotted-line) teams through resolution during critical deployment windows.

Customer-Facing Delivery: Serve as the primary technical point of contact for customers throughout planning, deployment, and validation, managing expectations and communicating status/issues clearly.

Partner Datacenter Coordination: Work with partner colocation facility teams to align on power, cooling, and space readiness ahead of and during deployments.

AI Factory Service Strategy: Contribute to strategies and improvements for Supermicro's end-to-end AI Factory deployment services, drawing on direct customer fulfillment experience to identify gaps and opportunities.

Documentation & Runbooks: Develop and maintain installation runbooks, deployment checklists, and troubleshooting playbooks to support repeatable, high-quality deployments as the team scales.

Qualifications:
  • Bachelor's degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Computer Science, or a related field-or equivalent experience-with 12+ years of experience in datacenter infrastructure or GPU cluster deployment.
  • Strong understanding of AI/NeoCloud infrastructure architectures and GPUaaS platforms.
  • Strong understanding of how software is built; able to direct and evaluate internal tooling work without necessarily writing production code.
  • Practical experience using Agentic AI tools and workflows for diagnostics or automation.
  • Hands-on experience planning, deploying, and validating GPU compute clusters (NVIDIA/CUDA preferred).
  • Working knowledge of BMS integration and datacenter infrastructure monitoring tools (e.g., Hyperview, Grafana, PagerDuty, or similar).
  • Solid understanding of datacenter cooling and power systems - liquid cooling (DLC), CDUs, Sidecars, cooling towers, chillers, and three-phase power distribution.
  • Excellent presentation and communication skills; comfortable and credible presenting technical recommendations directly to customers.
  • Demonstrated ability to lead cross-functional technical teams without formal authority (dotted-line leadership).
Preferred Qualifications:
  • Background in Electrical or Computer Engineering.
  • Prior experience deploying GPU clusters in colocation or multi-tenant partner datacenter environments.
  • Familiarity with high-density rack architectures (OCP, MGX, or similar).
  • Understanding of network fabrics used in GPU clusters (InfiniBand, RoCEv2).
  • Experience with ServiceNow CSM or similar customer-facing IT service platforms.
  • Experience building or directing agentic AI/coding workflows for infrastructure diagnostics or operations automation.
  • Interest or experience in contributing to service strategy or process improvement initiatives.
Salary Range:

$165,000 - $200,000

The salary offered will depend on several factors, including your location, level, education, training, specific skills, years of experience, and comparison to other employees already in this role. In addition to a comprehensive benefits package, candidates may be eligible for other forms of compensation, such as participation in bonus and equity award programs.

EEO Statement:

Supermicro is an Equal Opportunity Employer and embraces diversity in our employee population. It is the policy of Supermicro to provide equal opportunity to all qualified applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, protected veteran status or special disabled veteran, marital status, pregnancy, genetic information, or any other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Center Solutions Engineer
Staff Data Center Solutions Engineer

Support Revolution • San Jose (CA)

On-site
USD 165,000 - 200,000
Staff Data Center Solutions Engineer
Staff Data Center Solutions Engineer

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 165,000 - 200,000
Staff Data Center Solutions Engineer (29832)
Staff Data Center Solutions Engineer (29832)

Supermicro • San Jose (CA)

On-site
USD 165,000 - 200,000
Comprehensive benefits
Bonus and equity programs
Sr. Manager - Data Center Development (28721)
Sr. Manager - Data Center Development (28721)

Supermicro • San Jose (CA)

On-site
USD 175,000 - 195,000
Staff Data Center Solution Manager
Staff Data Center Solution Manager

Support Revolution • San Jose (CA)

On-site
USD 190,000 - 213,000
System Engineer, GPU Server
System Engineer, GPU Server

Supermicro • Wayne (CA)

On-site
USD 90,000 - 110,000
Sr. System Engineer - Network
Sr. System Engineer - Network

Supermicro • San Jose (CA)

On-site
USD 137,000 - 156,000
Participation in bonus and equity award programs
Comprehensive benefits package
System Engineer
System Engineer

Supermicro • Wayne (CA)

On-site
USD 140,000 - 158,000
Staff Data Center Electrical Engineer (27960)
Staff Data Center Electrical Engineer (27960)

Supermicro • San Jose (CA)

On-site
USD 175,000 - 200,000
System Engineer, GPU Server
System Engineer, GPU Server

Support Revolution • San Jose (CA)

On-site
USD 90,000 - 110,000