Senior Staff SRE: Global Infra & Cloud Reliability

NVIDIA

Santa Clara (CA)

On-site

USD 200,000 - 322,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA seeks a Senior Staff Site Reliability Engineer to lead modernization of IT Compute Core Team and scale core services across on‑prem and cloud environments. You will design, deploy, and optimize global infrastructure components for reliability and performance.

You will partner with leaders and engineers to deliver scalable IT products, implement observability, and drive efficiency using cutting-edge technologies. This role emphasizes automation and capacity planning.

Qualifications

  • Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience.
  • 12+ years of proven experience in compute platform engineering with a focus on automation.
  • Experience in designing and deploying Containerization architectures and Distributed Systems Infrastructure
  • Proven experience evaluating existing application architectures and identify opportunities for containerization to improve scalability, reliability, and efficiency.
  • Strong analytical skills with the ability to define and track key performance metrics.
  • Experience in developing tools for data analysis and performance profiling, Development with Terraform, Config Management tools.
  • Proficiency in programming languages such as Go and/or Python.
  • Linux OS Proficiency with Kernel Internals
  • Experience with running large environments consisting of BareMetal Build Infrastructure
  • Understanding of Network Protocols and Architectures (VLAN/VxLAN/SDN/BGP/Anycast)

Responsibilities

  • Lead initiatives to transform IT Compute Core Team, architecture to build new service offerings across On-Prem and Cloud
  • Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP.
  • Define and implement metrics to measure the efficiency of services and drive efficiency with software and hardware optimizations (SR-IOV/ DPU)
  • Experience with Technologies like eBPF and XDP for Observability & DDoS mitigation
  • Collect and review system data for capacity and planning purposes, analyze capacity data and develop plans for appropriate level enterprise-wide systems, and coordinate with management personnel in implementing changes.
  • Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, monitoring.
  • Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling IT products and services that meet customer needs.

Skills

Automation
Containerization
Distributed Systems
Metrics & Monitoring
Terraform
Configuration Management
Go/Python
Linux Kernel
BareMetal
Networking (VLAN/VxLAN/SDN/BGP/Anycast

Education

Bachelor's degree

Tools

eBPF
XDP

Job description

NVIDIA seeks a Senior Staff Site Reliability Engineer to lead modernization of IT Compute Core Team and scale core services across on‑prem and cloud environments. You will design, deploy, and optimize global infrastructure components for reliability and performance.

You will partner with leaders and engineers to deliver scalable IT products, implement observability, and drive efficiency using cutting-edge technologies. This role emphasizes automation and capacity planning.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff SRE: Global On-Prem & Cloud Infra Lead
Senior Staff SRE: Global On-Prem & Cloud Infra Lead

NVIDIA • United States

On-site
USD 200,000 - 322,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: Global HPC & Multi-Cloud Reliability
Senior SRE: Global HPC & Multi-Cloud Reliability

NVIDIA Corporation • Durham (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior SRE — Scale AI Systems, Equity Eligible
Senior SRE — Scale AI Systems, Equity Eligible

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale
Senior Site Reliability Engineer – AI-Driven, Hybrid Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE — Global HPC/AI Infra with Equity
Senior SRE — Global HPC/AI Infra with Equity

Socket.dev • North Carolina

On-site
USD 152,000 - 288,000
Senior SRE: Global HPC & Cloud Reliability
Senior SRE: Global HPC & Cloud Reliability

NVIDIA AI • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Benefits package
Lead Cloud SRE Architect for Private Cloud & AI CI/CD
Lead Cloud SRE Architect for Private Cloud & AI CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Site Reliability Lead — Global IT & Incident Command
Senior Site Reliability Lead — Global IT & Incident Command

NVIDIA Corporation • Durham (NC)

On-site
USD 184,000 - 265,000
Senior Site Reliability Lead — Enterprise IT & Automation
Senior Site Reliability Lead — Enterprise IT & Automation

NVIDIA Gruppe • Durham (NC)

On-site
USD 184,000 - 265,000
Equity
Benefits
On-site in Durham