Senior Site Reliability Engineering - Storage

NVIDIA AI

Bengaluru

On-site

INR 4,000,000 - 6,500,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA in Bengaluru, India seeks a Senior Site Reliability Engineer – Storage to own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power internal and external services.

You will combine storage expertise with automation and SRE practices to design, build, and operate highly available storage systems at scale, mentor engineers, and define SLOs/SLIs to drive continuous reliability improvements.

Qualifications

  • 12+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering with focus on storage systems.
  • Strong hands-on experience with NAS, SAN, and Object Storage platforms.
  • Solid understanding of SRE concepts (SLOs/SLIs, error budgets, incident management, observability, postmortems).
  • Proficiency with Infrastructure as Code and configuration tools (Terraform, Ansible, Puppet, SaltStack) and source control.
  • Experience building and operating highly available, scalable infrastructure with automation for provisioning, monitoring, and remediation.
  • Experience with containers and virtualization (Docker, Kubernetes) and modern CI/CD tools.
  • Strong scripting or programming skills (Python, Go, Shell).
  • Excellent communication and collaboration across distributed teams.
  • Bachelor’s degree in Computer Science/Engineering or related field (or equivalent).

Responsibilities

  • Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms.
  • Capture requirements from partner teams, architect storage solutions, and drive end-to-end implementation.
  • Develop and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management.
  • Participate in on-call and incident response, perform RCA and preventive actions.
  • Define and track SLOs/SLIs and use observability to improve reliability.
  • Build and maintain runbooks, SOPs, and documentation for storage services and automation.
  • Analyze capacity, perform forecasting, and recommend scaling strategies.
  • Collaborate with SRE, infrastructure, networking, and application teams; mentor junior engineers.
  • Share best practices and drive adoption of SRE principles.

Skills

SRE concepts
Observability
Automation
Incident management
Communication

Education

Bachelor's degree in Computer Science/Engineering

Tools

Terraform
Ansible
Puppet
SaltStack
Docker
Kubernetes
CI/CD
Git

Job description

Job Requisition ID

JR2024297

Job Category

Engineering

Time Type

Full time

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale.

What You Will Be Doing
  • Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security.
  • Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services.
  • Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure.
  • Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions.
  • Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuously improve reliability and efficiency.
  • Build and maintain runbooks, standard operating procedures, and comprehensive documentation for storage services and automation.
  • Analyze capacity and usage trends, perform forecasting, and recommend scaling or optimization strategies to support business growth.
  • Collaborate closely with SRE, infrastructure, networking, and application teams in a follow‑the‑sun model to deliver consistent, high‑quality service.
  • Mentor junior engineers, share best practices, and help drive adoption of SRE principles across the team.
What We Need To See
  • 12+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering, with significant focus on storage systems.
  • Strong hands‑on experience with design, deployment, and operations of enterprise‑grade NAS, SAN, and/or Object Storage platforms.
  • Solid understanding of SRE concepts (SLOs/SLIs, error budgets, incident management, observability, postmortems).
  • Proficiency with Infrastructure as Code and configuration management tools (e.g., Terraform, Ansible, Puppet, SaltStack) and source control systems.
  • Experience building and operating highly available, scalable infrastructure, including automation for provisioning, monitoring, and remediation.
  • Experience with container and virtualization platforms (e.g., Docker, Kubernetes, hypervisors) and modern CI/CD and version control tools.
  • Strong scripting or programming skills (e.g., Python, Go, Shell) to build tools, automate workflows, and integrate systems.
  • Excellent communication and collaboration skills, with the ability to work effectively across distributed and cross‑functional teams.
  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field (or equivalent practical experience).
Ways To Stand Out From The Crowd
  • Experience with storage for high‑performance computing, AI/ML workloads, or large‑scale data analytics.
  • Proven ability to debug complex, distributed systems and storage performance issues.
  • History of driving reliability improvements through data‑driven analysis and automation.
  • Experience leading technical initiatives, mentoring engineers, or acting as a technical lead on critical projects.

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most brilliant and talented people in the world working for us. If you're creative and motivated, we want to hear from you!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineering, Storage
Senior Site Reliability Engineering, Storage

NVIDIA Corporation • India

On-site
INR 4,500,000 - 7,000,000
Senior High-Performance Storage Architect - NVIS
Senior High-Performance Storage Architect - NVIS

NVIDIA Corporation • India

On-site
INR 14,231,000 - 19,924,000
Senior High-Performance Storage Architect – NVIS
Senior High-Performance Storage Architect – NVIS

NVIDIA Corporation • Mumbai

On-site
INR 17,013,000 - 22,684,000
Senior High-Performance Storage Architect - NVIS
Senior High-Performance Storage Architect - NVIS

NVIDIA Gruppe • Mumbai

On-site
INR 13,233,000 - 19,849,000
Senior High-Performance Storage Architect - NVIS
Senior High-Performance Storage Architect - NVIS

NVIDIA GraphicsPLtd,Pune • India

On-site
INR 17,078,000 - 22,770,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 4,000,000 - 6,500,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 5,000,000 - 7,500,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA Gruppe • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA Corporation • India

On-site
INR 3,500,000 - 6,000,000