Datacenter Hardware & Linux Systems Engineer (GPU Clusters)

Sciforium

San Jose (CA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity

Job summary

Sciforium is seeking a Hardware Operations & Systems Engineer to oversee the physical health of GPU clusters in San Jose, California. This role includes managing system health, deploying hardware, and maintaining Linux operating systems to ensure that teams have a stable environment for compute workloads.

Ideal candidates will have experience in Linux Systems Administration, server hardware troubleshooting, and networking security. Benefits include medical insurance, a 401k plan, and flexible time off.

Qualifications

  • 3+ years of experience in Linux Systems Administration, including boot processes and disk management.
  • Strong background in server hardware troubleshooting, particularly in high-density environments.
  • Experience managing networking security with VPNs and directory services like LDAP.
  • Proficiency in Bash for system automation.

Responsibilities

  • Serve as the primary contact for physical system outages and hardware failures.
  • Monitor hardware health, including GPU thermals and power draw.
  • Coordinate with data center staff and third-party vendors for repairs and maintenance.
  • Install, patch, and maintain Linux operating systems across servers.

Skills

Linux Systems Administration
Server Hardware Troubleshooting
Networking Security Management
Bash Scripting

Tools

Ansible
Ubuntu
CentOS
RHEL

Job description

Sciforium is seeking a Hardware Operations & Systems Engineer to oversee the physical health of GPU clusters in San Jose, California. This role includes managing system health, deploying hardware, and maintaining Linux operating systems to ensure that teams have a stable environment for compute workloads.

Ideal candidates will have experience in Linux Systems Administration, server hardware troubleshooting, and networking security. Benefits include medical insurance, a 401k plan, and flexible time off.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cluster Ops Engineer: Linux & Hardware
GPU Cluster Ops Engineer: Linux & Hardware

Sciforium • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
GPU Cluster Engineer, Hardware Operations
GPU Cluster Engineer, Hardware Operations

Sciforium • San Francisco (CA)

On-site
USD 120,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+4
Datacenter Field Engineer
Datacenter Field Engineer

Sciforium • San Jose (CA)

On-site
USD 100,000 - 130,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Networking
GPU Cluster Engineer, Networking

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
GPU Systems Engineer for AI Training Clusters
GPU Systems Engineer for AI Training Clusters

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1