Senior SRE Ops Lead — On‑Site, Equity Eligible

NVIDIA

Durham (NC)

On-site

USD 184,000 - 265,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA seeks a Site Reliability Operations Technical Lead in Durham, NC to own and elevate reliability for the site. You will supervise incidents, drive root-cause analysis, and coordinate with global teams while mentoring site engineers and maintaining runbooks.

Responsibilities include patching, security hardening, automation development, and executive-level communication during outages. This role requires on-site work, after-hours support, and a rigorous, hands-on approach.

Qualifications

  • 12+ years in enterprise support engineering, infrastructure, or end user services, including 5+ years in a senior, lead, or escalation-tier role in a multi-site environment.
  • Deep hands-on experience across Active Directory and hybrid Entra ID, Exchange hybrid, Windows and Linux server, virtualization, enterprise storage, and datacenter hardware.
  • Enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf, or equivalent), Windows 11, and the Microsoft 365 ecosystem, plus endpoint security and vulnerability remediation.
  • Database operations support and networking fundamentals — DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting.
  • Scripting and automation in Python, PowerShell, or Bash applied to real support problems, and ServiceNow or similar ITSM.
  • Demonstrated technical leadership without formal authority, executive-level communication during incidents, and root-cause focus.
  • Willingness to work on-site and hands-on (lifting/moving equipment), join on-call rotation, and after-hours maintenance windows.
  • Bachelor's degree in Computer Science, Information Systems, or related field, or equivalent experience.

Responsibilities

  • Own day-to-day site operations — incidents, requests, critical issues, and support coverage; manage queue health, SLA attainment, backlog, and service quality.
  • Serve as Tier 3 escalation owner for the site and AMER across identity, messaging, compute, and endpoint domains; drive root cause and permanent fixes.
  • Own endpoint compliance, vulnerability remediation, patch management, and hardening; maintain audit readiness and support with InfoSec on incidents.
  • Drive issues into global platform teams and vendors with reproduction cases and diagnostic evidence to a committed fix.
  • Act as technical lead for site SRO engineers; set standards, review work, mentor, and manage the knowledge base/runbooks.
  • Represent site IT in incidents, changes, onboarding/moves, and office/lab expansions with leadership alignment.
  • Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks, and reporting; pursue AI-driven, proactive solutions.
  • Represent site and AMER priorities in regional/global IT initiatives, architecture forums, and operating-rhythm decisions.

Skills

Active Directory
Entra ID
Exchange
Windows Server
Linux Server
Virtualization
Storage
Datacenter hardware
Intune
Autopilot
MECM/SCCM
Jamf
Scripting
PowerShell
Python
Bash
ServiceNow
ITSM
Incident management
Executive communication
On‑call

Education

Bachelor's degree in Computer Science/Information Systems

Tools

PowerShell
Python
Bash
ServiceNow
ITSM
Jamf

Job description

NVIDIA seeks a Site Reliability Operations Technical Lead in Durham, NC to own and elevate reliability for the site. You will supervise incidents, drive root-cause analysis, and coordinate with global teams while mentoring site engineers and maintaining runbooks.

Responsibilities include patching, security hardening, automation development, and executive-level communication during outages. This role requires on-site work, after-hours support, and a rigorous, hands-on approach.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Ops Lead — On-Site, Equity Eligible
Senior SRE Ops Lead — On-Site, Equity Eligible

Relha LLC • Durham (NC), Northern (KY)

Hybrid
USD 184,000 - 265,000
Equity
Benefits
Senior Site Reliability Lead — Enterprise IT & Automation
Senior Site Reliability Lead — Enterprise IT & Automation

NVIDIA Gruppe • Durham (NC)

On-site
USD 184,000 - 265,000
Equity
Benefits
On-site in Durham
Senior Site Reliability Lead — Global IT & Incident Command
Senior Site Reliability Lead — Global IT & Incident Command

NVIDIA Corporation • Durham (NC)

On-site
USD 184,000 - 265,000
Senior SRE Lead, Site Reliability Operations (Equity)
Senior SRE Lead, Site Reliability Operations (Equity)

NVIDIA Gruppe • Seattle (WA)

On-site
USD 184,000 - 265,000
Equity
Benefits
Senior Site Reliability Lead - On-Site Ops & Automation
Senior Site Reliability Lead - On-Site Ops & Automation

Relha LLC • Seattle (WA), Northern (KY)

Hybrid
USD 184,000 - 265,000
Equity
Senior Site Reliability Lead – Enterprise IT Ops
Senior Site Reliability Lead – Enterprise IT Ops

NVIDIA Corporation • Seattle (WA)

On-site
USD 184,000 - 265,000
Equity
Benefits
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Operations Lead (Seattle)
Senior Site Reliability Operations Lead (Seattle)

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 265,000
Senior SRE — Scale AI Systems, Equity Eligible
Senior SRE — Scale AI Systems, Equity Eligible

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Staff SRE: Global Infra & Cloud Reliability
Senior Staff SRE: Global Infra & Cloud Reliability

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits