Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 272,000 - 431,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team. You will design and support end-to-end CI/CD for GPU development, onboard internal teams, and optimize AI workflows with scalable, cost-efficient cloud services.

As a senior technical leader, you will guide a team of engineers, tackle complex distributed systems, and build robust metrics. This role demands deep expertise in cloud infra, AI provisioning, and cross-organizational collaboration

Qualifications

  • BS or MS in Electrical Engineering, Computer Science, or related field.
  • 15+ years of systems software development incl. AI exposure.
  • Experience in maintaining cloud infrastructure and highly available production environments.
  • Strong programming skills in Java, Python, and shell scripting with distributed systems knowledge.
  • Experience with SQL/NoSQL databases such as MySQL, Cassandra, MongoDB or Elasticsearch.
  • Proficient with Docker containers and virtual machines.
  • Knowledge of cloud technologies like OpenStack, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.
  • Ability to collaborate across teams and time zones in a multinational environment.

Responsibilities

  • Serve an SRE Architect on the GPU Private Cloud team used by NVIDIA engineers for development and CI/CD.
  • Evaluate and develop software solutions to optimize workflows across NVIDIA groups.
  • Architect and support end-to-end CI/CD systems using open-source and NVIDIA tools.
  • Onboard internal development teams to Private Cloud with use-case discovery and solution mapping.
  • Identify bottlenecks and optimize performance and cost of AI development and testing systems.
  • Lead software projects and guide engineers to deliver impactful solutions.
  • Diagnose problems in software systems and resolve issues.
  • Create metrics and dashboards using analytics to monitor cloud performance.

Skills

Java
Python
Shell scripting
Distributed systems
REST APIs
MySQL
NoSQL
Docker
Virtual machines
OpenStack
Kubernetes
Git

Education

BS/MS in Electrical Engineering or Computer Science

Tools

OpenStack
Kubernetes
Chef/Puppet
Hadoop/Ceph/SwiftStack
Git

Job description

NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. These cloud services provide almost half a million automated jobs per day on thousands of servers helping with the efficiency of thousands of NVIDIA's software engineers worldwide. The cloud hosts various machines and devices with operating systems like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers unified CI/CD solutions and cloud-based software development. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?

What you'll be doing:
  • Serve an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing.
  • Evaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.
  • Architecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software.
  • Customer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud.
  • Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.
  • Leading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions.
  • Looking for problems within software systems and resolving the issues
  • Craft and implement critical metrics using various analytics methods and dashboards.
What we need to see:
  • BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).
  • 15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.
  • Experience of maintaining cloud infrastructure and highly available production environment.
  • Strong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.
  • Experience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch.
  • Excellent knowledge and working experience with Docker containers and Virtual Machines.
  • Good background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.
  • Ability to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment.
Ways to stand out from the crowd:
  • Depth in AI, Machine Learning and Deep Learning algorithms and techniques.
  • Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.
  • Experience developing large-scale software systems using modular architecture under real-time performance requirements.
  • Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 9, 2026.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Solutions Architect, IPP
Senior Solutions Architect, IPP

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity options
Comprehensive benefits package
Diverse work environment
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Solutions Architect, IPP
Senior Solutions Architect, IPP

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Comprehensive benefits package
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity compensation
Benefits package
Competitive salary
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

NVIDIA • United States

On-site
USD 320,000 - 489,000
Senior Infrastructure Solutions Architect
Senior Infrastructure Solutions Architect

NVIDIA Corporation • Austin (TX)

On-site
USD 180,000 - 290,000
Principal Software Engineer - Cloud Services
Principal Software Engineer - Cloud Services

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000