GPU Data Center Operations Engineer

Cadence Design Systems

San Jose (CA)

On-site

USD 119,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cadence Design Systems in California is seeking a Data Center Operations Engineer to support, maintain, and deploy critical data center infrastructure with a focus on Linux-based systems, GPU server deployments, and InfiniBand networking. This role collaborates with global infrastructure, development, and operations teams to ensure reliable service delivery.

The position requires hands-on experience with data center operations, cluster bring-up, hardware installation, and troubleshooting across

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
  • Strong hands-on experience in Linux environments, including system administration, troubleshooting, and performance validation.
  • Proficiency with Linux command-line tools and shell scripting (Bash or equivalent).
  • Experience with cluster bring-up, driver installation, and system-level configuration.
  • Hands-on experience setting up and validating GPU servers in clustered environments.
  • Experience with end-to-end GPU testing in InfiniBand-based clusters.
  • Working knowledge of InfiniBand networking, including switch configuration and subnet management.
  • Solid understanding of networking fundamentals, including the OSI model and TCP/IP protocol suite (IP, ARP, ICMP, TCP, UDP, SMTP, FTP, TFTP).
  • Experience installing, configuring, and troubleshooting routers, switches, and terminal servers.
  • Familiarity with fiber and copper cabling, including IP and SAN deployments.
  • Experience managing incident tickets, maintaining acceptable ticket loads, and meeting SLAs.
  • Strong organizational skills with meticulous attention to detail in data center environments.
  • Ability to follow and enforce documented escalation procedures and operational policies.
  • Strong verbal and written communication skills, with the ability to collaborate effectively with cross-functional and global teams.

Responsibilities

  • Provide hands-on operational support for all data center projects, deployments, and repair activities.
  • Participate in an on-call rotation and provide on-site or remote support during maintenance windows and incidents.
  • Troubleshoot and resolve operational issues related to Linux servers, GPU platforms, networking, and storage infrastructure.
  • Support customer and internal deployments, ensuring timely and successful bring-up of GPU servers and clusters.
  • Perform InfiniBand fabric bring-up, switch configuration, subnet management, and troubleshooting.
  • Conduct daily health checks of Linux systems and infrastructure components, proactively identifying and mitigating risks.
  • Install, configure, test, and maintain server hardware (rack and stack, labeling, HDDs, memory, CPUs, RAID batteries, NICs, etc.).
  • Install, configure, and troubleshoot networking equipment including routers, switches, and terminal servers for out-of-band management.
  • Review and validate equipment deployments against approved design documentation and standards.
  • Support data center builds, refreshes, migrations, and expansions while adhering to quality and safety standards.
  • Coordinate with vendors and onsite staff for hardware delivery, diagnostics, replacement, and warranty services.
  • Utilize monitoring and alerting frameworks to identify issues, elevate appropriately, and ensure timely service restoration.
  • Maintain accurate documentation of operational procedures, system configurations, and runbooks.
  • Follow established incident management, escalation procedures, and service-level agreements (SLAs).
  • Collaborate with global teams across time zones to support operational initiatives and continuous improvement efforts.
  • Contribute to process improvement initiatives and ensure adherence to documented policies, processes, and procedures.

Skills

Linux
GPU servers
InfiniBand
Cluster bring-up
Networking
Storage
Troubleshooting
On-call

Education

Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience

Job description

Cadence Design Systems in California is seeking a Data Center Operations Engineer to support, maintain, and deploy critical data center infrastructure with a focus on Linux-based systems, GPU server deployments, and InfiniBand networking. This role collaborates with global infrastructure, development, and operations teams to ensure reliable service delivery.

The position requires hands-on experience with data center operations, cluster bring-up, hardware installation, and troubleshooting across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Data Center Systems Engineer
GPU Data Center Systems Engineer

Cadence Design Systems, Inc. • San Jose (CA)

On-site
USD 78,523 - 146,025
GPU Data Center Operations Engineer
GPU Data Center Operations Engineer

Cadence • San Jose (CA)

On-site
USD 120,000 - 221,000
401(k) match
Employee stock purchase plan
Medical/dental/vision coverage
+1
Data Center Operations Engineer
Data Center Operations Engineer

Cadence • San Jose (CA)

On-site
USD 120,000 - 221,000
401(k) match
Employee stock purchase plan
Medical/dental/vision coverage
+1
Data Center Operations Engineer
Data Center Operations Engineer

Cadence Design Systems • San Jose (CA)

On-site
USD <1,000
GPU Data Center Engineer — Hybrid/Remote
GPU Data Center Engineer — Hybrid/Remote

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Data Center Operations Engineer
Data Center Operations Engineer

Cadence Design Systems, Inc. • San Jose (CA)

On-site
USD 78,523 - 146,025
GPU Data Center Deployment Specialist
GPU Data Center Deployment Specialist

Cirrascale Corporation • Austin (TX)

Hybrid
Health insurance
Dental insurance
Vision insurance
+2
Senior Data Center Network Engineer – GPU Clusters
Senior Data Center Network Engineer – GPU Clusters

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation, including meaningful equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+4
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2