Sr Software Engineer

Nvidia

Bengaluru

On-site

INR 4,500,000 - 7,500,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nvidia in Bengaluru is seeking a Senior Software Engineer to design, build, and operate scalable platforms for observability, automation, and AI-driven reliability across enterprise infrastructure. You will own large-scale distributed services spanning storage, compute, network, VMware, OpenShift, and bare-metal, delivering proactive operations, telemetry pipelines, and autonomous remediation.

Collaborate across teams, mentor engineers, and influence architecture while staying focused on

Qualifications

  • 10+ years of software engineering, SRE, or distributed-systems experience with technical leadership.
  • Strong Go or Python production-grade distributed systems experience.
  • Proven ownership of complex software/platform initiatives across multiple teams.
  • Deep understanding of distributed systems, event-driven architectures and high-throughput data processing.
  • Experience with modern observability and telemetry platforms.
  • SRE and infrastructure knowledge across Kubernetes/OpenShift, VMware, bare-metal, storage, networking, and cloud environments.
  • Ability to influence direction, mentor engineers, and deliver measurable outcomes.

Responsibilities

  • Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at scale.
  • Develop reusable platform services, APIs, automation frameworks, and control planes for self-service and reduced toil.
  • Build telemetry and event-processing systems spanning metrics, logs, traces, events, and alerts at billions of signals.
  • Develop AI-native reliability capabilities including anomaly detection and automated remediation.
  • Drive technical architecture across Storage, Compute, Network, and Platform domains.
  • Engineer for production at scale with a focus on quality, security, performance, and maintainability.
  • Provide technical leadership and mentorship to influence engineering standards and architecture decisions.

Skills

Go
Python
Distributed systems
SRE
Observability
AI
Platform leadership

Education

Bachelor's or Master's in Computer Science or equivalent

Tools

Kafka
NATS
gRPC
Terraform
Ansible
OpenTelemetry
Prometheus
VictoriaMetrics
Vector
Loki
Grafana
ClickHouse
FastAPI

Job description

We are seeking a Senior Software Engineer with strong infrastructure expertise to design, build, and operate the next generation of our enterprise Observability, Automation, and AI-driven Reliability Platform.

This role will build highly scalable distributed systems and platform services spanning Storage, Compute, Network, VMware, OpenShift, and bare-metal infrastructure. The engineer will help transform infrastructure operations from reactive monitoring and manual remediation to proactive, predictive, and AI-driven autonomous operations.

What You Will Be Doing:
  • Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at large scale.
  • Develop reusable platform services, APIs, automation frameworks, and control planes that enable self-service, reduce operational toil, and automate infrastructure operations across multiple engineering teams.
  • Build scalable telemetry and event-processing systems spanning metrics, logs, traces, events, topology, and alerts, with the performance and efficiency to process billions of infrastructure signals.
  • Build intelligent and AI-native reliability capabilities, including agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation.
  • Drive technical architecture and engineering direction across Storage, Compute, Network, and Platform domains, solving complex and ambiguous problems that span multiple teams.
  • Engineer for production at scale, with strong focus on software quality, scalability, security, performance, observability, maintainability, and operational readiness.
  • Provide technical leadership and mentorship, influence engineering standards and architecture decisions, and deliver measurable improvements in reliability, MTTR, operational toil, engineering productivity, and infrastructure efficiency.
What We Need To See:
  • Bachelors or Masters degree in Computer Science, Engineering, or equivalent practical experience, with 10+ years of software engineering, SRE, infrastructure, or distributed-systems experience and demonstrated technical leadership.
  • Strong software engineering expertise in Go, Python, or equivalent languages, with experience designing and building production-grade distributed systems, platform services, APIs, and automation.
  • Proven experience owning complex software/platform initiatives across multiple teams or infrastructure domains, from architecture and implementation through adoption and measurable impact.
  • Deep understanding of distributed systems, event-driven architectures, microservices, APIs, and high-throughput data processing, including technologies such as Kafka, NATS, gRPC, or equivalent.
  • Strong experience with modern observability and telemetry platforms, including OpenTelemetry, Prometheus, VictoriaMetrics, Vector, Loki, Grafana, ClickHouse, or equivalent technologies.
  • Strong SRE and infrastructure knowledge across Kubernetes/OpenShift, VMware, bare-metal, storage, networking, and/or cloud environments, with experience using Terraform, Ansible, or equivalent automation technologies.
  • Demonstrated ability to solve ambiguous problems, influence technical direction without direct authority, mentor engineers, establish engineering standards, and deliver measurable operational and business outcomes.
Ways To Stand Out From The Crowd:
  • Experience building software and reliability platforms for large-scale on-premises infrastructure, particularly Storage, Compute, Networking, VMware, and Kubernetes/OpenShift.
  • Deep understanding of storage and infrastructure telemetry, including IOPS, latency, NVMe health, SAN/NAS topology, block/object storage, and infrastructure failure domains.
  • Experience building self-healing systems, automated remediation, predictive operations, or autonomous SRE capabilities.
  • Production experience applying Generative AI, AIOps, LLMs, or Agentic AI to incident triage, RCA, operational intelligence, debugging, or remediation; experience with LangChain, LlamaIndex, AutoGen, or equivalent is a plus.
  • Experience building high-performance platform services using FastAPI, gRPC, or equivalent technologies.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer
Senior Software Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA Corporation • India

On-site
INR 3,500,000 - 6,000,000
Senior Software Engineer
Senior Software Engineer

NVIDIA AI • Bengaluru

On-site
INR 4,000,000 - 6,000,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Senior Staff SRE – Compute Platform
Senior Staff SRE – Compute Platform

NVIDIA Gruppe • Bengaluru

On-site
INR 400,000 - 700,000
Senior Staff SRE – Compute Platform
Senior Staff SRE – Compute Platform

NVIDIA • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Senior Devops Engineer
Senior Devops Engineer

Bounteous • Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,400,000