Senior Staff Site Reliability Engineer

NVIDIA Gruppe

Bengaluru

On-site

INR 3,500,000 - 7,000,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA invites an experienced Senior Site Reliability Engineer to lead the technical strategy and roadmap for large-scale SRE initiatives across enterprise systems. You will design resilient distributed systems powering NVIDIA's AI-enabled enterprise products and services, transforming legacy apps into scalable architectures.

You will drive automation and observability, apply AI workload signals, and develop autonomous incident response pipelines.

Qualifications

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
  • BS degree in Computer Science or related field involving coding, or equivalent experience.
  • Strong proficiency in Python, Typescript, JavaScript, or Go with automation and IaC focus.
  • Experience with IaC tooling (AWS CDK, CloudFormation, Terraform, CrossPlane).
  • Solid understanding of OpenTelemetry or observability at scale, incl. AI workloads.
  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS/Azure/GCP).
  • Outstanding problem-solving, communication, and collaboration skills.

Responsibilities

  • Lead the technical strategy and roadmap for large-scale SRE initiatives to boost reliability and scalability.
  • Design and build resilient distributed systems powering AI-enabled enterprise products.
  • Drive automation and observability improvements using metrics and AI signals.
  • Build LLM-aware monitoring and autonomous incident response pipelines to reduce toil and MTTR.
  • Collaborate with Cloud, Platform, Security, and AI/ML teams to ensure high availability and secure ops.
  • Analyze and run complex systems, including Kubernetes-scale and AI/ML infra challenges.
  • Promote AI-first engineering practices and mentor engineers in agentic development workflows.

Skills

Python
Typescript
JavaScript
Go
Automation
Distributed systems debugging
Kubernetes
Public cloud

Education

BS in Computer Science or related field

Tools

AWS CDK
Terraform
CloudFormation
CrossPlane
OpenTelemetry
Kubernetes
Azure
GCP

Job description

What you\'ll be doing:


  • Lead the technical strategy and roadmap for large-scale, multi-functional SRE initiatives that boost reliability, scalability, and developer efficiency throughout enterprise systems.

  • Design and build resilient distributed systems that power NVIDIA\'s next-generation AI-powered enterprise products and services, including transforming legacy applications and database systems into modern and scalable architectures.

  • Drive automation and observability improvements, using metrics and analytics — including AI workload quality signals and model performance telemetry — to improve performance, reliability, and efficiency.

  • Build LLM-aware monitoring and autonomous incident response pipelines to reduce toil, accelerate MTTR, and evolve on-call operations toward AI-assisted remediation.

  • Work together with Cloud, Platform, Security, and AI/ML groups to develop modern SRE elements and AI-native platform features that guarantee high availability and secure operations.

  • Analyze and run complex systems — including Kubernetes-scale and AI/ML infrastructure challenges — championing standards in system design and incident management.

  • Drive AI-assisted and AI-first engineering practices across the organization, mentoring engineers to adopt agentic development workflows, coding agents, and LLM-powered tooling in their day-to-day work.


What we need to see:


  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.

  • BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience.

  • Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code and distributed systems debugging.

  • Experience with infrastructure-as-code tooling such as AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.

  • Solid understanding of OpenTelemetry or other observability implementations at scale, including observability build for AI workloads (model performance, drift detection, and AI quality signals).

  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP), including bare-metal and GPU-accelerated infrastructure.

  • Outstanding problem-solving, communication, and collaboration skills, with the ability to influence across technical and interpersonal boundaries.


Ways to stand out from the crowd:


  • Experience with Public Cloud or large-scale automation systems.

  • Capability to steer technical strategy and achieve quantifiable reliability results in complex, multi-team settings.

  • Experience building or operating agentic AI platforms — autonomous or semi-autonomous — including LLM toolchains, agent orchestration frameworks (e.g., LangGraph, AutoGen), or automated runbook and toil-reduction processes.

  • A strong sense of ownership, curiosity, and innovation — and turn challenges into opportunities.


NVIDIA leads the charge in innovative breakthroughs in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, functions as the visual cortex of today\'s computers and forms the core of our products and services. Our work opens new realms to explore, encourages outstanding creativity and discovery, and powers inventions once thought of as science fiction — from artificial intelligence to autonomous systems. NVIDIA is searching for outstanding talent like you to help us advance the next wave of artificial intelligence!


Widely considered to be one of the technology world\u2019s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 4,000,000 - 6,500,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 5,000,000 - 7,500,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 1,500,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA Corporation • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 900,000 - 1,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NVIDIA • New Delhi

On-site
INR 4,000,000 - 8,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 300,000 - 550,000
Senior Solutions Architect - Generative AI
Senior Solutions Architect - Generative AI

NVIDIA Gruppe • Pune District

On-site
INR 2,000,000 - 3,500,000
Comprehensive benefits package
Highly competitive salaries
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 4,000,000 - 7,000,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA • Pune District

On-site
INR 1,500,000 - 2,100,000