Senior Datacenter Resiliency Architect

TieTalent

Santa Clara (CA)

On-site

USD 130,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A technology company is seeking a Senior Datacenter Resiliency Architect to enhance the resiliency features of their industry-leading Datacenter GPUs. Responsibilities include architecting resiliency solutions, performing simulations, and collaborating with diverse teams. The ideal candidate has a Master’s or PhD in a related field, along with 5+ years of relevant experience in GPU architecture. Strong skills in Python and C/C++ are essential. The position is based in Santa Clara, California.

Qualifications

  • 5+ years of relevant experience in GPU and networking architectures.
  • Strong knowledge and experience in RAS features.
  • Familiarity with computer architecture basics and machine learning concepts.

Responsibilities

  • Architect hardware and software resiliency features.
  • Collaborate to align verification requirements.
  • Develop and implement architecture verification test plans.

Skills

GPU hardware architecture
Python scripting
C/C++ programming
Architectural debugging
Resiliency features
Interpersonal skills
Collaboration

Education

Master’s or PhD in Computer Engineering or Electrical Engineering

Job description

Join to apply for the Senior Datacenter Resiliency Architect role at TieTalent

We are seeking a Senior Datacenter Resiliency (RAS) Architect to support the development and validation of GPU hardware and software resiliency features. You will be a key member of a team of innovators, challenging the status quo and pushing beyond boundaries, with impact on the industry’s leading Datacenter GPUs and SOCs powering AI and HPC products.

What you’ll be doing
  • Architect hardware and software resiliency features to improve system Reliability, Availability, Serviceability (RAS), and performance in the Datacenter.
  • Model and analyze RAS metrics (e.g., Failures in Time for permanent and transient errors, Availability from GPU to Rack to Datacenter); use models to identify gaps and drive RAS improvements.
  • Collaborate with architects, unit designers, and software engineers to ensure alignment of verification requirements.
  • Develop and implement comprehensive architecture verification test plans for resiliency features.
  • Execute Architecture Test Plan by developing test content and enabling, running, and debugging tests on architecture models; support test debug on RTL, emulation, and silicon.
  • Run simulations to analyze Architectural Vulnerability Factor and liveness of on-die memory, flip-flops, and latches.
  • Develop CUDA software diagnostics kernels to run on clusters of NVIDIA GPUs to identify hardware issues.
  • Develop and automate fault models to simulate various fault types (e.g., transient faults, stuck-at faults) in gate-level netlists, RTL, architectural models, silicon, and other environments.
What we need to see
  • Master’s or PhD in Computer Engineering, Electrical Engineering, or closely related field, or equivalent experience.
  • At least 5+ years of relevant experience.
  • Familiarity with GPU and networking architectures, computer architecture basics (caches, coherence, buses, DMA), and machine learning/deep learning concepts.
  • Strong knowledge and experience in GPU hardware architecture or RAS features, or both.
  • Proficiency in developing architecture models.
  • Scripting and automation with Python or similar; proficiency in C/C++.
  • Excellent interpersonal skills and ability to collaborate with on-site and remote teams; strong debugging and analytical skills; self-driven and results oriented.
  • Experience with resiliency and datacenter RAS or Verilog/SystemVerilog RTL simulations and debugging; ability to set up test benches and integrate components is a plus.
  • Programming with CUDA is a plus.
Company/role notes

NVIDIA’s work spans high-performance computing and AI computing—roles involve building resilient, high-availability computing platforms for AI, HPC, and data center workloads. NVIDIA is an equal opportunity employer; we do not discriminate on protected characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Datacenter Resiliency Architect for GPUs
Senior Datacenter Resiliency Architect for GPUs

TieTalent • Santa Clara (CA)

On-site
USD 130,000 - 160,000
Senior SoC Architect, RAS
Senior SoC Architect, RAS

NVIDIA AI • Eugene (OR)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior SoC Architect, RAS
Senior SoC Architect, RAS

NVIDIA Corporation • Town of Texas (WI), Northern (KY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior SoC Architect, RAS
Senior SoC Architect, RAS

NVIDIA • Hillsboro (OR)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior SoC Architect, RAS
Senior SoC Architect, RAS

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior SoC Architect, RAS
Senior SoC Architect, RAS

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior SoC Architect, RAS
Senior SoC Architect, RAS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 230,000 - 340,000
Equity compensation
Benefits package
Distinguished Resiliency and Safety Architect, GPU Diagnostics
Distinguished Resiliency and Safety Architect, GPU Diagnostics

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior SoC RAS Architect for Resilient AI Platforms
Senior SoC RAS Architect for Resilient AI Platforms

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior SoC RAS Architect – Drive Resilient AI Compute
Senior SoC RAS Architect – Drive Resilient AI Compute

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000