Senior Observability Platform Engineer: Scale & Automation

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

NVIDIA in Santa Clara, CA is seeking a Senior Systems Software Engineer for the Observability and Telemetry Platform to design, build, and operate large-scale production systems. You’ll combine software and systems engineering to ensure reliability, uptime, and performance while enabling developers to ship changes safely.

You’ll work across Linux, networking, containers, and cloud tooling (Kubernetes, OpenStack, Docker) to reduce manual work, automate routines, and improve scalability and

Qualifications

  • BS degree in Computer Science or a related technical field.
  • 5+ years of infrastructure automation experience.
  • 5+ years delivering infrastructure and observability platforms.
  • Experience with Linux, networking and containers.

Responsibilities

  • Design, implement and support operational and reliability aspects of large-scale Observability & Telemetry platform.
  • Engage in and improve the whole lifecycle of services—from design through deployment and refinement.
  • Support services before they go live through design consulting, tooling, capacity management and launch reviews.
  • Maintain services once live by measuring availability, latency and health.
  • Scale systems via automation and improvements to reliability and velocity.

Skills

Infrastructure automation
Distributed systems
Observability platforms
Python
Go
Perl
Ruby
Linux
Networking
Containers

Education

BS in Computer Science or related field

Tools

Kubernetes
OpenStack
Docker
Grafana
OpenTelemetry
Prometheus

Job description

NVIDIA in Santa Clara, CA is seeking a Senior Systems Software Engineer for the Observability and Telemetry Platform to design, build, and operate large-scale production systems. You’ll combine software and systems engineering to ensure reliability, uptime, and performance while enabling developers to ship changes safely.

You’ll work across Linux, networking, containers, and cloud tooling (Kubernetes, OpenStack, Docker) to reduce manual work, automate routines, and improve scalability and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer - Scale Reliability
Senior Observability Platform Engineer - Scale Reliability

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Observability & Telemetry Platform Engineer
Senior Observability & Telemetry Platform Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Systems Automation Engineer, Large-Scale Infra
Senior Systems Automation Engineer, Large-Scale Infra

NVIDIA Corporation • Durham (NC)

On-site
USD 190,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior AI/HPC Telemetry & Observability Engineer
Senior AI/HPC Telemetry & Observability Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Senior AIOps & Observability Architect
Senior AIOps & Observability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior GPU Platform Engineer — Systems, Observability & AI
Senior GPU Platform Engineer — Systems, Observability & AI

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000