Senior Staff Site Reliability Engineer – Compute Platform

NVIDIA Gruppe

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA seeks a Senior Staff SRE to build and operate reliable, scalable compute platforms supporting global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations.

You will lead provisioning and lifecycle management in data centers, develop automation, observability tooling, and define SLOs/SLIs, incident response, and blameless postmortems.

Qualifications

  • BS in Computer Science, Engineering, or related field, or equivalent experience.
  • 10+ years operating production infrastructure or platform services.
  • Strong Kubernetes, KubeVirt, Docker, and Linux systems expertise.
  • Experience provisioning bare-metal infrastructure and data-center operations.
  • Proficiency in Python or Go and building RESTful services with infrastructure APIs.
  • Experience with Terraform and configuration management tools (Ansible, Chef, Puppet).
  • Strong observability knowledge with OpenTelemetry, Prometheus, Grafana, and logging stacks.
  • Clear written and interpersonal communication; demonstrated delivery of scalable fixes.

Responsibilities

  • Build, operate, and improve large-scale compute platforms with focus on performance, reliability, and scale.
  • Lead bare-metal provisioning and lifecycle management in data centers (PXE, DHCP, DNS, OS provisioning).
  • Develop automation, self-service capabilities, and observability solutions via APIs and IaC.
  • Define and operate SLOs/SLIs; lead incident investigations and blameless postmortems.
  • Partner with cross-functional teams to deliver global platform initiatives and participate in on-call rotations.

Skills

Kubernetes administration
KubeVirt
Docker
Containerization
Linux systems
Programming (Python or Go)
Infrastructure as Code
Observability
SRE concepts (SLI/SLO/_error budgets)
Incident management

Education

BS in Computer Science or related field

Tools

Terraform
Ansible
Chef/Puppet
OpenTelemetry
Prometheus
Grafana
ELK Stack
Splunk

Job description

NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services.

What you’ll be doing:
  • Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale.
  • Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.
  • Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data.
  • Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems.
  • Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation.
What we need to see:
  • BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services.
  • Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges.
  • Experience deploying and operating bare-metal infrastructure in a data-center environment, including provisioning, networking, operating-system lifecycle management, and hardware automation.
  • <
  • Proficiency in Python, Go, or a comparable programming language, with experience building RESTful services and integrating infrastructure APIs.
  • Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, Chef, or Puppet, along with a solid understanding of TCP/IP networking and infrastructure security.
  • Strong SRE and observability experience, including SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.
  • Clear written and interpersonal communication skills, with a record of delivering practical, scalable solutions to complex technical problems.
Ways to stand out from the crowd:
  • Experience operating HPC, AI, GPU-accelerated, or general-purpose bare-metal compute infrastructure, including GPU-enabled Kubernetes or KubeVirt clusters.
  • Expertise with VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.
  • Experience applying generative AI or agentic workflows to improve infrastructure diagnostics, reduce operational toil, and accelerate incident resolution.
  • Experience building secure, integrated operational platforms using APIs, RBAC, service accounts, secrets management, audit controls, workflow orchestration, and infrastructure or incident-management systems.
  • Demonstrated delivery of complex, high-impact infrastructure projects.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Staff SRE – Compute Platform
Senior Staff SRE – Compute Platform

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 700,000
Senior Staff SRE – Compute Platform
Senior Staff SRE – Compute Platform

NVIDIA • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Staff Site Reliability Engineer – Compute Platform
Senior Staff Site Reliability Engineer – Compute Platform

NVIDIA • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Staff Site Reliability Engineer – Compute Platform
Senior Staff Site Reliability Engineer – Compute Platform

NVIDIA AI • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Senior Staff SRE – Compute Platform
Senior Staff SRE – Compute Platform

NVIDIA Corporation • India

On-site
INR 4,500,000 - 9,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA Corporation • India

On-site
INR 3,500,000 - 6,000,000
Senior Site Reliability Engineer, Production Engineering
Senior Site Reliability Engineer, Production Engineering

NVIDIA • India

Hybrid
INR 2,500,000 - 4,500,000
Senior Site Reliability Engineer, Production Engineering
Senior Site Reliability Engineer, Production Engineering

NVIDIA • Bengaluru

On-site
INR 2,800,000 - 5,200,000
Senior Software Engineer
Senior Software Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000