Senior Site Reliability Engineer

NVIDIA

New Delhi

On-site

INR 4,000,000 - 8,000,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a Senior Site Reliability Engineer to join its cloud service team and support the GeForce NOW platform. The role focuses on reliability, scalability, and automation across cloud and datacenter environments.

You will lead incident response, drive observability initiatives, and collaborate with multiple teams to improve service resilience and customer experience. A BS degree and 5+ years in production SRE are required.

Qualifications

  • BS degree in Computer Science, Computer Engineering, Information Technology, or a related technical field (or equivalent).
  • 5+ years of experience supporting and operating production services in a live-site environment as an SRE/Production Engineer.
  • Strong understanding of containerization, microservices, and Kubernetes and related ecosystem components.
  • Troubleshoot complex production issues and drive resolution.
  • Experience with observability platforms and cloud services (AWS, Azure, GCP).
  • Hands-on automation experience with Python, Go, Bash, or similar languages.

Responsibilities

  • Monitor, support, and maintain reliability, availability, and performance of GeForce NOW production services across cloud and data center environments.
  • Participate in production incident triage, troubleshooting, and resolution; participate in on-call rotation including weekend coverage.
  • Monitor service health with metrics, logs, traces, and dashboards; identify reliability and capacity issues before impact.
  • Collaborate with software engineering, platform, networking, and infrastructure teams to improve readiness and resilience.
  • Drive observability improvements by enhancing monitoring, alerting, dashboards, and telemetry.
  • Scale services by building automation and improving deployment, recovery, and operational workflows.
  • Lead incident response, root cause analysis, and blameless post-mortems; drive corrective actions.
  • Design and develop custom tools and self-service solutions to simplify operations and improve productivity.
  • Evaluate processes and identify opportunities to improve reliability and customer experience.
  • Contribute to deployment and operation of Kubernetes-based services for scalability and performance.

Skills

Kubernetes
Python
Go
Bash
SLOs/SLIs
Incident management

Education

BS degree in CS/CE/IT or related field

Tools

Prometheus
Grafana
ELK/OpenSearch
OpenTelemetry

Job description

We are now looking for a Sr. Site Reliability Engineer (SRE)! NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s motivated by outstanding technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing – an era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. NVIDIA is at the forefront of generative AI models, from language to images. Doing what’s never been done before takes vision, innovation, and the world’s best talent, and as an NVIDIA you’ll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work.

NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its cloud service team to support, triage, and build the GeForce NOW cloud gaming platform. As SREs, you will manage the big picture of how our systems relate to each other using a breadth of tools and approaches to tackle a broad spectrum of problems. We live SRE practices that are key to product quality – limiting time spent on reactive operational work, conducting blameless post‑mortems, proactively identifying potential outages, and iteratively improving our services. The role includes responsibility for service response and workflow, driving tools and service development to maintain and improve SLOs, partnering with service owners to enhance reliability, and continuously evaluating operational processes to identify opportunities for improvement via custom tools, automation, and self‑service solutions that improve overall reliability and operational efficiency of GeForce NOW.

What You Will Be Doing
  • Monitor, support, and maintain the reliability, availability, and performance of large‑scale GeForce NOW production services running across cloud and datacenter environments.
  • Participate in production incident triage, troubleshooting, and resolution of complex infrastructure and application issues; take part in the team’s on‑call rotation, including occasional weekend coverage, to ensure timely restoration of customer‑facing services.
  • Monitor service health using metrics, logs, traces, and dashboards, and proactively identify reliability, performance, and capacity issues before they impact customers.
  • Collaborate with software engineering, platform, networking, and infrastructure teams to improve operational readiness, reliability, and service resilience.
  • Drive observability initiatives by improving monitoring, alerting, dashboards, and telemetry to enable faster detection and diagnosis of production issues.
  • Scale services sustainably by building automation, eliminating operational toil, and continuously improving deployment, recovery, and operational workflows.
  • Lead and participate in incident response, root cause analysis, and blameless post‑mortems, driving corrective and preventive actions to improve long‑term service reliability.
  • Design and develop custom tools, automation, and self‑service solutions that simplify operations, improve engineer productivity, and enhance the overall GeForce NOW platform.
  • Continuously evaluate existing operational processes and identify opportunities to improve service reliability, operational efficiency, and customer experience through engineering‑driven solutions.
  • Contribute to the design, deployment, and operation of Kubernetes‑based services, ensuring they meet scalability, reliability, and performance requirements.
What We Need To See
  • BS degree in Computer Science, Computer Engineering, Information Technology, or a related technical field (or equivalent experience).
  • 5+ years of experience supporting and operating mission‑critical production services in a live‑site environment as a Site Reliability Engineer (SRE), Production Engineer, or similar role.
  • Strong understanding of containerization, microservices architecture, and Kubernetes, including Kubernetes ecosystem components and operational best practices.
  • Demonstrated ability to troubleshoot complex production issues, identify root causes, and drive issues to resolution.
  • Strong understanding of distributed systems and how complex production environments interact across applications, infrastructure, networking, and cloud services.
  • Experience supporting production operations, including incident management, change management, postmortem reviews, and operational excellence initiatives.
  • Hands‑on experience developing automation using Python, Go, Bash, or similar scripting/programming languages.
  • Strong understanding of SLOs, SLIs, error budgets, KPIs, and service reliability best practices.
  • Experience with observability platforms such as Prometheus, Grafana, ELK/OpenSearch, and modern monitoring and alerting solutions.
  • Experience operating services in public cloud environments such as AWS, Azure, GCP, or equivalent cloud platforms.
Ways To Stand Out From The Crowd
  • Experience supporting large‑scale customer‑facing cloud or gaming services.
  • Strong Kubernetes operational and troubleshooting expertise.
  • Experience with observability platforms, including Prometheus, Grafana, ELK/OpenSearch, and OpenTelemetry.
  • Strong scripting or programming skills in Python, Go, or similar languages with a focus on automation.
  • Experience driving production incident response, postmortems, and operational excellence initiatives.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people in the world working for us. If you’re creative and autonomous, we want to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 300,000 - 550,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 4,000,000 - 7,000,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 5,000,000 - 7,500,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 4,000,000 - 6,500,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA Corporation • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Site Reliability Engineer
Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 1,500,000 - 2,100,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior DevOps Engineer
Senior DevOps Engineer

NVIDIA • Pune District

On-site
INR 1,500,000 - 2,100,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000
Senior Cloud Software Engineer
Senior Cloud Software Engineer

NVIDIA • Bengaluru

On-site
INR 2,000,000 - 3,000,000