Lead Cloud SRE Architect for Private Cloud & AI CI/CD

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 272,000 - 431,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team. You will design and support end-to-end CI/CD for GPU development, onboard internal teams, and optimize AI workflows with scalable, cost-efficient cloud services.

As a senior technical leader, you will guide a team of engineers, tackle complex distributed systems, and build robust metrics. This role demands deep expertise in cloud infra, AI provisioning, and cross-organizational collaboration

Qualifications

  • BS or MS in Electrical Engineering, Computer Science, or related field.
  • 15+ years of systems software development incl. AI exposure.
  • Experience in maintaining cloud infrastructure and highly available production environments.
  • Strong programming skills in Java, Python, and shell scripting with distributed systems knowledge.
  • Experience with SQL/NoSQL databases such as MySQL, Cassandra, MongoDB or Elasticsearch.
  • Proficient with Docker containers and virtual machines.
  • Knowledge of cloud technologies like OpenStack, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.
  • Ability to collaborate across teams and time zones in a multinational environment.

Responsibilities

  • Serve an SRE Architect on the GPU Private Cloud team used by NVIDIA engineers for development and CI/CD.
  • Evaluate and develop software solutions to optimize workflows across NVIDIA groups.
  • Architect and support end-to-end CI/CD systems using open-source and NVIDIA tools.
  • Onboard internal development teams to Private Cloud with use-case discovery and solution mapping.
  • Identify bottlenecks and optimize performance and cost of AI development and testing systems.
  • Lead software projects and guide engineers to deliver impactful solutions.
  • Diagnose problems in software systems and resolve issues.
  • Create metrics and dashboards using analytics to monitor cloud performance.

Skills

Java
Python
Shell scripting
Distributed systems
REST APIs
MySQL
NoSQL
Docker
Virtual machines
OpenStack
Kubernetes
Git

Education

BS/MS in Electrical Engineering or Computer Science

Tools

OpenStack
Kubernetes
Chef/Puppet
Hadoop/Ceph/SwiftStack
Git

Job description

NVIDIA is seeking a Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team. You will design and support end-to-end CI/CD for GPU development, onboard internal teams, and optimize AI workflows with scalable, cost-efficient cloud services.

As a senior technical leader, you will guide a team of engineers, tackle complex distributed systems, and build robust metrics. This role demands deep expertise in cloud infra, AI provisioning, and cross-organizational collaboration

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI-Driven Cloud SRE Architect for Private Cloud & CI/CD
AI-Driven Cloud SRE Architect for Private Cloud & CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Cloud SRE Architect - AI/CI-CD & Equity Options
Cloud SRE Architect - AI/CI-CD & Equity Options

NVIDIA AI • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Rack-Scale AI Cloud Architect — Innovative Data Center Solutions
Rack-Scale AI Cloud Architect — Innovative Data Center Solutions

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 414,000
Equity
Benefits
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
AI Infra Architect — GPU HPC & Cloud/On‑Prem
AI Infra Architect — GPU HPC & Cloud/On‑Prem

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA AI • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Cloud Solutions Architect for AI & CI/CD
Senior Cloud Solutions Architect for AI & CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity options
Comprehensive benefits package
Diverse work environment
Senior AI Frameworks Engineer — CI/CD for Multi-Cloud
Senior AI Frameworks Engineer — CI/CD for Multi-Cloud

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits