Site Reliability Engineer

Sustainable Talent

Santa Clara (CA)

On-site

USD 89,544 - 117,096

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Full benefits
Paid time off
Company culture experience

Job summary

Sustainable Talent is partnering with Nvidia in Santa Clara for a Site Reliability Engineer (Contract) position. This full-time role offers pay ranging from $65/hr to $85/hr based on various factors and includes full benefits and paid time off.

The ideal candidate will have a strong background in maintaining and setting up Linux and Windows hosts, excellent debugging skills, and experience in scripting. This role involves collaborating within the cloud team to support infrastructure needs effectively.

Qualifications

  • 5+ years of experience in large-scale enterprise production systems.
  • Experience in debugging infrastructure issues.
  • Proficiency in scripting with Python or Go.

Responsibilities

  • Monitor and recover assets in the private cloud environment.
  • Deploy and maintain a large farm of machines.
  • Contribute to monitoring systems for real-time insights.

Skills

Debugging and analytical skills
Python or Go scripting
Linux and Windows systems maintenance

Education

Bachelor’s or Master’s Degree in Computer Science or Software Engineering

Tools

Chef
Ansible
Terraform
Version control systems such as Perforce or Git

Job description

Overview

Sustainable Talent is partnering with Nvidia, a global leader in computer graphics, PC gaming, and accelerated computing, to bring a Site Reliability Engineer (Contract) to the team based in Santa Clara, CA.

This is a full‑time (W‑2) contract role. Pay ranges from $65/hr to $85/hr based on experience, education, location, and other factors. The position includes full benefits, paid time off, and a year‑long company culture experience.

The Cloud team is an infrastructure group within NVIDIA’s IPP (Infrastructure, Planning and Process). It collaborates with NVIDIA Software groups such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence, and Driverless Cars to support their infrastructure needs. The cloud hosts nearly half a million automated jobs per day on thousands of servers, enabling efficiency for NVIDIA’s software engineers worldwide.

The cloud environment hosts a heterogeneous mix of machines and devices across Windows, Linux, and Android operating systems, with a wide range of NVIDIA GPU and Tegra processor hardware. The SRE will build the next generation of cloud services, craft solutions, mine data to uncover real problems, and resolve them.

Responsibilities
  • Fleet monitoring and recovery of assets in the private cloud environment that houses compute servers with NVIDIA GPUs.
  • Build and stabilize virtualization infrastructure using ESXi, KVM, and Hyper‑V.
  • Deploy and maintain a large farm of machines with configuration management and infrastructure automation tools (Chef, Ansible, Terraform).
  • Participate in on‑call and rotational L1 support for continuous monitoring and remediation of infrastructure issues (PagerDuty).
  • Analyze and debug operating system, networking, configuration, and performance problems.
  • Assist in rollout and deployment of infrastructure configurations to support the latest NVIDIA hardware and technologies.
  • Contribute to the development of monitoring systems to provide fast, reliable real‑time insight into infrastructure subsystems (Zabbix, Big Panda, Grafana).
Qualifications
  • Bachelor’s or Master’s Degree in Computer Science, Software Engineering, or equivalent experience.
  • Five or more years of professional experience working in large‑scale enterprise production systems.
  • Strong debugging and analytical skills to triage, root‑cause, and resolve infrastructure issues in collaboration with the platform engineering team.
  • Experience maintaining and setting up Linux and Windows hosts.
  • Proficiency in scripting with Python or Go; Unix shell experience.
  • Experience with version control systems such as Perforce or Git.
Ways to Stand Out
  • Experience with virtualization and hardware virtualization technologies such as VMware, KVM, Hyper‑V, Docker, and Kubernetes.
  • Background in automating bare‑metal and VM provisioning.
  • Experience supporting GPUs, embedded device development, driver development, and CUDA/TensorRT applications.
  • Development experience with Chef, Ansible, and infrastructure orchestration.
Equal Employment Opportunity Statement

Sustainable Talent is a M/F+, disabled, and veteran equal employment opportunity and affirmative action employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Site Reliability Engineer - Cloud
Distinguished Site Reliability Engineer - Cloud

2100 NVIDIA USA • United States

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Highly competitive salaries
Comprehensive benefits package
Equity eligibility
Senior Site Reliability Engineer - HPC
Senior Site Reliability Engineer - HPC

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Equity
Comprehensive benefits package
Site Reliability Engineer - Hardware Infrastructure
Site Reliability Engineer - Hardware Infrastructure

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Comprehensive benefits package
Service Reliability Engineer
Service Reliability Engineer

NVIDIA Gruppe • Town of Texas (WI)

On-site
USD 168,000 - 334,000
Service Reliability Engineer
Service Reliability Engineer

Nvidia Corporation in • Austin (TX)

On-site
USD 208,000 - 334,000
Equity
Benefits
Service Reliability Engineer
Service Reliability Engineer

NVIDIA AI • Town of Texas (WI)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer, GeForce NOW
Senior Site Reliability Engineer, GeForce NOW

NVIDIA • California (MO)

On-site
USD 168,000 - 270,000
Equity
Benefits
Senior Site Reliability Engineer - Cloud
Senior Site Reliability Engineer - Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 265,000
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits