Senior AIOps & Observability Platform Architect

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 200,000 - 322,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Senior Software Engineer to design and develop AIOps and Observability platforms used across cloud, on-prem, and data centers. You will lead the architecture, implement metrics, logs, traces, and dashboards, and guide teams on observability best practices.

Expect to mentor engineers and collaborate with data scientists on ML-driven anomaly detection. You will drive the technical vision, align with business goals, and build scalable, secure, and reliable systems capable of

Qualifications

  • Bachelor's degree in computer science and engineering or related field.
  • 12+ years in product development and full-stack engineering with 5+ years in observability platforms.
  • Strong knowledge of Prometheus, Victoria Metrics, Grafana, OpenTelemetry, and related tooling.
  • Hands-on experience with Kubernetes, Docker, and microservices architectures.
  • Experience with streaming systems handling large-scale events.

Responsibilities

  • Lead design, development, and deployment of AIOps and Observability platforms.
  • Define roadmap and standards for observability across NVIDIA.
  • Collaborate with teams to address observability needs and ensure scalable solutions.
  • Provide peer reviews focusing on performance, security, and reliability.
  • Work with data scientists on ML models for anomaly detection and root-cause analysis.

Skills

Go
Python
Java
C#
Mentoring

Education

Bachelor's degree in CS/Engineering

Tools

Prometheus
Victoria Metrics
Vector
Loki
Grafana
OpenTelemetry
Clickhouse
PagerDuty
Datadog
Kubernetes
Docker
Kafka

Job description

NVIDIA is seeking a Senior Software Engineer to design and develop AIOps and Observability platforms used across cloud, on-prem, and data centers. You will lead the architecture, implement metrics, logs, traces, and dashboards, and guide teams on observability best practices.

Expect to mentor engineers and collaborate with data scientists on ML-driven anomaly detection. You will drive the technical vision, align with business goals, and build scalable, secure, and reliable systems capable of

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AIOps & Observability Architect
Senior AIOps & Observability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior AIOps & Observability Platform Engineer
Senior AIOps & Observability Platform Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior AI Observability Platform Engineer
Senior AI Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA AI • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior Observability Platform Engineer & AIOps Lead
Senior Observability Platform Engineer & AIOps Lead

Software Technology Inc. • Buffalo (NY)

On-site
USD 140,000 - 200,000
Senior Observability Platform Engineer, GPU AI Cloud
Senior Observability Platform Engineer, GPU AI Cloud

nscaleoperationsukltd • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000