AI-Driven Cloud SRE Architect for Private Cloud & CI/CD

NVIDIA

Santa Clara (CA)

On-site

USD 272,000 - 431,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

NVIDIA is seeking an experienced Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team in Santa Clara. The role focuses on building and scaling private cloud infrastructure, CI/CD pipelines, and AI development environments for NVIDIA engineers worldwide.

Key responsibilities include leading projects, optimizing workflows, and ensuring high availability. A strong background in distributed systems, AI-related development, and cloud technologies is essential for

Qualifications

  • BS/MS in Electrical Engineering, Computer Science, or equivalent experience.
  • 15+ years of systems software development with AI exposure.
  • Experience in maintaining cloud infrastructure and highly available production environments.
  • Strong programming and software development skills in Java, Python, and shell scripting.
  • Solid understanding of distributed systems and REST APIs.
  • Experience with SQL/NoSQL databases (MySQL, Cassandra, MongoDB, Elasticsearch).
  • Excellent knowledge of Docker and virtual machines.
  • Proficiency with cloud tech like OpenStack, Docker, Kubernetes, and related tools.

Responsibilities

  • Serve as an SRE Architect for the GPU Private Cloud used by NVIDIA teams globally.
  • Identify and develop software solutions to optimize critical development workflows.
  • Architect, implement, and support end-to-end CI/CD systems using open-source and NVIDIA tooling.
  • Onboard internal NVIDIA teams to Private Cloud with use-case discovery and solution mapping.
  • Identify performance bottlenecks and optimize AI development and testing environments.
  • Lead software projects and guide engineers to deliver impactful solutions.
  • Craft metrics and dashboards to monitor cloud performance.

Skills

Java
Python
Shell scripting
REST APIs
Distributed systems

Education

BS/MS in Electrical Engineering or Computer Science

Tools

OpenStack
Docker
Kubernetes
Chef/Puppet
Hadoop/Ceph/SwiftStack
LXC
Git
Perforce
JFrog
Kafka

Job description

NVIDIA is seeking an experienced Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team in Santa Clara. The role focuses on building and scaling private cloud infrastructure, CI/CD pipelines, and AI development environments for NVIDIA engineers worldwide.

Key responsibilities include leading projects, optimizing workflows, and ensuring high availability. A strong background in distributed systems, AI-related development, and cloud technologies is essential for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Cloud SRE Architect for Private Cloud & AI CI/CD
Lead Cloud SRE Architect for Private Cloud & AI CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Cloud SRE Architect – AI/ML CI/CD Leader (Equity)
Senior Cloud SRE Architect – AI/ML CI/CD Leader (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Cloud SRE Architect - AI/CI-CD & Equity Options
Cloud SRE Architect - AI/CI-CD & Equity Options

NVIDIA AI • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior AI Frameworks Engineer — CI/CD for Multi-Cloud
Senior AI Frameworks Engineer — CI/CD for Multi-Cloud

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Rack-Scale AI Cloud Architect — Innovative Data Center Solutions
Rack-Scale AI Cloud Architect — Innovative Data Center Solutions

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 414,000
Equity
Benefits
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Manager, AI-Powered Site Reliability
Senior Manager, AI-Powered Site Reliability

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 248,000 - 397,000
Equity and benefits