Cloud SRE Architect - AI/CI-CD & Equity Options

NVIDIA AI

Santa Clara (CA)

On-site

USD 272,000 - 431,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team. You will lead SRE efforts for the GPU Private Cloud, optimize AI development workflows, and architect end-to-end CI/CD solutions across NVIDIA.

The role requires extensive experience in distributed systems, cloud platforms, and leadership of engineering teams, with a base salary range of 272000–431250 USD and equity eligibility.

Qualifications

  • BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent).
  • 15+ years of systems software development with AI experience.
  • Experience maintaining cloud infrastructure and highly available production environments.
  • Strong programming in JAVA, Python, Shell and understanding of distributed systems and REST APIs.

Responsibilities

  • Serve as an SRE Architect for GPU Private Cloud used by NVIDIA developers worldwide.
  • Develop software solutions to optimize AI development workflows across NVIDIA teams.
  • Architect, implement, and support end-to-end CI/CD systems using open-source and NVIDIA tools.
  • Onboard internal teams to Private Cloud with case discovery and solution mapping.
  • Identify bottlenecks and optimize speed and cost of AI development and testing systems.
  • Lead software projects and guide engineers to deliver impactful solutions.
  • Craft metrics and dashboards using analytics methods.

Skills

Distributed systems
REST APIs
Python
Java
Shell scripting
Leadership
AI/ML awareness

Education

BS or MS in Electrical Engineering or Computer Science

Tools

Docker
Virtual Machines
OpenStack
Kubernetes
Chef/Puppet
Git
Perforce
JFrog
Kafka
MySQL
Cassandra
MongoDB
Elasticsearch

Job description

NVIDIA is seeking a Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team. You will lead SRE efforts for the GPU Private Cloud, optimize AI development workflows, and architect end-to-end CI/CD solutions across NVIDIA.

The role requires extensive experience in distributed systems, cloud platforms, and leadership of engineering teams, with a base salary range of 272000–431250 USD and equity eligibility.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Cloud SRE Architect for Private Cloud & AI CI/CD
Lead Cloud SRE Architect for Private Cloud & AI CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
AI-Driven Cloud SRE Architect for Private Cloud & CI/CD
AI-Driven Cloud SRE Architect for Private Cloud & CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA AI • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Cloud Solutions Architect for AI & CI/CD
Senior Cloud Solutions Architect for AI & CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity options
Comprehensive benefits package
Diverse work environment
Principal Engineer, Cloud Site Reliability Engineering
Principal Engineer, Cloud Site Reliability Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior Cloud Infrastructure Engineer — Scalable AI Ops
Senior Cloud Infrastructure Engineer — Scalable AI Ops

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Site Reliability Engineer
Site Reliability Engineer

Sustainable Talent • Santa Clara (CA)

On-site
Full benefits
Paid time off
Company culture experience