Senior Observability & Telemetry Platform Engineer

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation is seeking a Senior Systems Software Engineer for the Observability and Telemetry Platform. You will design, build, and maintain large-scale production systems that power GPU cloud services with high reliability, performance, and real-time monitoring.

The role requires expertise across software and systems engineering, Linux, networking, containers, and modern observability tools such as Grafana, OpenTelemetry, and Prometheus.

Qualifications

  • BS degree in Computer Science or a related technical field, or equivalent experience.
  • 5+ years in infrastructure automation and distributed systems designing production-scale platforms.
  • Experience delivering foundational infrastructure and observability platforms.
  • Proficiency with Python, Go, Perl or Ruby.
  • Strong knowledge of Linux, networking and container technologies.
  • Interest in large-scale distributed systems and problem solving.
  • Strong communication and ownership mindset.

Responsibilities

  • Design, implement and support operational and reliability aspects of large-scale Observability & Telemetry platform with focus on performance at scale.
  • Engage in and improve the entire service lifecycle from design to deployment and refinement.
  • Support services before go-live through design reviews, tooling, capacity management and launch reviews.
  • Maintain services in production by monitoring availability, latency and health.
  • Scale systems via automation and drive changes to improve reliability and velocity.
  • Practice blameless postmortems and incident response.
  • Participate in on-call rotation to support production systems.

Skills

Infrastructure automation
Distributed systems design
Python/Go/Perl/Ruby
Linux systems
Networking & Containers
Observability tooling

Education

BS in Computer Science or related field

Tools

Kubernetes
OpenStack
Docker
Grafana
OpenTelemetry
Prometheus

Job description

NVIDIA Corporation is seeking a Senior Systems Software Engineer for the Observability and Telemetry Platform. You will design, build, and maintain large-scale production systems that power GPU cloud services with high reliability, performance, and real-time monitoring.

The role requires expertise across software and systems engineering, Linux, networking, containers, and modern observability tools such as Grafana, OpenTelemetry, and Prometheus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior AIOps & Observability Platform Architect
Senior AIOps & Observability Platform Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior AIOps & Observability Architect
Senior AIOps & Observability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior AI Infrastructure Engineer Observability & Automation
Senior AI Infrastructure Engineer Observability & Automation

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 356,000
Equity
Benefits
Senior Observability Engineer: Telemetry for GPU Cloud
Senior Observability Engineer: Telemetry for GPU Cloud

Submer - Datacenters That Make Sense • United States

Remote
USD 104,000 - 174,000
Hybrid-friendly approach
International team diversity
Flexible work environment
Senior Software Engineer: Fleet Intelligence & GPU Telemetry
Senior Software Engineer: Fleet Intelligence & GPU Telemetry

NVIDIA • New York (NY)

On-site
USD 152,000 - 242,000
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

NVIDIA Corporation • Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000