Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 272,000 - 431,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation is seeking a Principal Engineer for Cloud Site Reliability Engineering in Santa Clara, CA. You will architect and lead GPU Private Cloud infrastructure supporting thousands of NVIDIA developers, with emphasis on scalable CI/CD and AI workloads.

The role requires 15+ years in systems software, strong Java/Python/shell scripting skills, and deep experience with cloud tech (Docker, VMs, OpenStack, Kubernetes). This full-time position offers equity and competitive benefits.

Qualifications

  • BS or MS in Electrical Engineering, Computer Science, or a relevant field (or equivalent experience).
  • 15+ years of systems software development experience, including at least 1 year focusing on AI exploration/development.
  • Experience maintaining cloud infrastructure and highly available production environments.
  • Strong programming skills in Java, Python, and shell scripting with a solid grasp of distributed systems and REST APIs.
  • Experience with SQL/NoSQL databases (MySQL, Cassandra, MongoDB, Elasticsearch).
  • Excellent knowledge of Docker containers and Virtual Machines.

Responsibilities

  • Serve as an SRE Architect for the GPU Private Cloud used by NVIDIA teams worldwide.
  • Evaluate and develop software to optimize core development workflows.
  • Architect, implement, and support end-to-end CI/CD systems with open-source and NVIDIA tooling.
  • Onboard internal development teams to Private Cloud with use-case discovery and solution mapping.
  • Identify bottlenecks and optimize performance and cost for AI development and testing pipelines.
  • Lead software projects and guide engineers toward efficient solutions.

Skills

Java
Python
Shell scripting
Distributed systems
REST APIs

Education

BS in Electrical Engineering, Computer Science, or relevant field

Tools

MySQL
Cassandra
MongoDB
Elasticsearch
Docker
OpenStack
Kubernetes
Chef/Puppet
Hadoop
Ceph
SwiftStack
LXC
Git
Perforce
JFrog
Kafka
Virtual Machines

Job description

NVIDIA Corporation is seeking a Principal Engineer for Cloud Site Reliability Engineering in Santa Clara, CA. You will architect and lead GPU Private Cloud infrastructure supporting thousands of NVIDIA developers, with emphasis on scalable CI/CD and AI workloads.

The role requires 15+ years in systems software, strong Java/Python/shell scripting skills, and deep experience with cloud tech (Docker, VMs, OpenStack, Kubernetes). This full-time position offers equity and competitive benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud SRE Architect – AI/ML CI/CD Leader (Equity)
Senior Cloud SRE Architect – AI/ML CI/CD Leader (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Lead Cloud SRE Architect for Private Cloud & AI CI/CD
Lead Cloud SRE Architect for Private Cloud & AI CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
AI-Driven Cloud SRE Architect for Private Cloud & CI/CD
AI-Driven Cloud SRE Architect for Private Cloud & CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Senior Cloud Solutions Architect for AI/CI/CD
Senior Cloud Solutions Architect for AI/CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Comprehensive benefits package
Senior AI Frameworks Engineer — CI/CD for Multi-Cloud
Senior AI Frameworks Engineer — CI/CD for Multi-Cloud

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior SRE — GPU Cloud, Large-Scale Clusters & Kubernetes
Senior SRE — GPU Cloud, Large-Scale Clusters & Kubernetes

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity