Senior Observability Platform Engineer - Scale Reliability

NVIDIA

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Senior Systems Software Engineer to design, build and maintain a large-scale Observability and Telemetry platform. The role combines software and systems engineering to ensure high reliability, uptime and fuel developer productivity through automation and performance tuning.

You will work on end-to-end service lifecycles, from design to deployment, and contribute to capacity planning, monitoring and incident response in a diverse, collaborative environment that values bold

Qualifications

  • BS degree in Computer Science or a related technical field
  • 5+ years of experience with infrastructure automation and distributed systems design
  • 5+ years delivering foundational infrastructure and observability platforms
  • Experience in Python, Go, Perl or Ruby
  • In-depth knowledge of Linux, Networking and Containers

Responsibilities

  • Design, implement and support operational and reliability aspects of large-scale Observability & Telemetry platform
  • Engage in and improve the full lifecycle of services from inception to deployment
  • Support services before go-live via system design consulting and capacity management
  • Maintain services post-launch by monitoring availability and latency
  • Scale systems through automation and reliability improvements
  • Conduct blameless postmortems and on-call rotations

Skills

Infrastructure automation
Distributed systems
Python
Go
Linux
Networking
Containers

Education

BS in Computer Science or related field

Tools

Kubernetes
OpenStack
Docker
Grafana
OpenTelemetry
Prometheus

Job description

NVIDIA is seeking a Senior Systems Software Engineer to design, build and maintain a large-scale Observability and Telemetry platform. The role combines software and systems engineering to ensure high reliability, uptime and fuel developer productivity through automation and performance tuning.

You will work on end-to-end service lifecycles, from design to deployment, and contribute to capacity planning, monitoring and incident response in a diverse, collaborative environment that values bold

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer: Scale & Automation
Senior Observability Platform Engineer: Scale & Automation

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Observability & Telemetry Platform Engineer
Senior Observability & Telemetry Platform Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AIOps & Observability Platform Architect
Senior AIOps & Observability Platform Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior AIOps & Observability Architect
Senior AIOps & Observability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior Systems Software Engineer, Observability and Telemetry Platform
Senior Systems Software Engineer, Observability and Telemetry Platform

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Systems Automation Engineer, Large-Scale Infra
Senior Systems Automation Engineer, Large-Scale Infra

NVIDIA Corporation • Durham (NC)

On-site
USD 190,000 - 357,000
Equity
Benefits