GPU Data Center Operations Engineer

Cadence

San Jose (CA)

On-site

USD 120,000 - 221,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) match
Employee stock purchase plan
Medical/dental/vision coverage
Paid vacation

Job summary

Cadence is seeking a Data Center Operations Engineer to support, maintain, and deploy critical data center infrastructure with a focus on Linux systems, GPU deployments, and InfiniBand networks. You’ll work with global teams to ensure reliable, secure, and scalable service delivery across compute, network, and storage environments.

You will perform hardware bring‑up, server and GPU validation, and incident response while adhering to SLAs and safety standards.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • Strong hands‑on Linux administration, troubleshooting, and performance validation.
  • Proficiency with Linux CLI and shell scripting (Bash).
  • Experience with cluster bring‑up, driver installation, and system configuration.
  • Hands‑on experience setting up GPU servers in clusters.
  • Experience with end‑to‑end GPU testing in InfiniBand clusters.
  • Knowledge of InfiniBand networking, switches, and subnet management.
  • Solid networking fundamentals: TCP/IP, IP, ARP, ICMP, etc.

Responsibilities

  • Provide hands‑on operational support for data center projects, deployments, and repair activities.
  • Participate in on‑call rotation and provide on‑site or remote support during maintenance windows and incidents.
  • Troubleshoot Linux servers, GPU platforms, networking, and storage infrastructure.
  • Support customer and internal deployments, bringing up GPU servers and clusters.
  • Perform InfiniBand fabric bring‑up, switch configuration, and subnet management.
  • Conduct daily health checks of Linux systems and infrastructure components.
  • Install and maintain server hardware including rack, HDDs, memory, CPUs, RAID, NICs.
  • Install and troubleshoot routers, switches, and terminal servers for out‑of‑band management.
  • Review deployments against design docs and standards.
  • Support data center builds, migrations, and expansions per SLAs.
  • Coordinate with vendors and onsite staff for hardware delivery and repairs.
  • Use monitoring to identify issues and ensure timely service restoration.
  • Maintain runbooks and documentation of procedures and configurations.
  • Follow incident management and escalation procedures.
  • Collaborate with global teams across time zones on initiatives.
  • Contribute to process improvements and policy adherence.

Skills

Linux system administration
Shell scripting (Bash)
GPU server deployments
InfiniBand networking
Networking fundamentals (TCP/IP)
Incident management & SLAs
Documentation & runbooks
Cross-functional collaboration

Education

Bachelor’s degree in CS/Engineering/IT or equivalent

Tools

Linux CLI tools
InfiniBand switches
GPU drivers installation

Job description

Cadence is seeking a Data Center Operations Engineer to support, maintain, and deploy critical data center infrastructure with a focus on Linux systems, GPU deployments, and InfiniBand networks. You’ll work with global teams to ensure reliable, secure, and scalable service delivery across compute, network, and storage environments.

You will perform hardware bring‑up, server and GPU validation, and incident response while adhering to SLAs and safety standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Data Center Systems Engineer
GPU Data Center Systems Engineer

Cadence Design Systems, Inc. • San Jose (CA)

On-site
USD 78,523 - 146,025
GPU Data Center Operations Engineer
GPU Data Center Operations Engineer

Cadence Design Systems • San Jose (CA)

On-site
USD <1,000
Data Center Operations Engineer
Data Center Operations Engineer

Cadence • San Jose (CA)

On-site
USD 120,000 - 221,000
401(k) match
Employee stock purchase plan
Medical/dental/vision coverage
+1
Data Center Operations Engineer
Data Center Operations Engineer

Cadence Design Systems • San Jose (CA)

On-site
USD <1,000
GPU Data Center Engineer — Hybrid/Remote
GPU Data Center Engineer — Hybrid/Remote

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
Data Center Operations Engineer
Data Center Operations Engineer

Cadence Design Systems, Inc. • San Jose (CA)

On-site
USD 78,523 - 146,025
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits
Datacenter Operations Manager
Datacenter Operations Manager

Cadence • San Jose (CA)

On-site
USD 150,000 - 230,000