Senior AIOps & Observability Platform Engineer

NVIDIA AI

Santa Clara (CA)

On-site

USD 200,000 - 322,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a highly skilled Senior Software Engineer to design and develop AIOps and Observability platforms. You will mentor engineers and collaborate with product managers to define strategy and roadmaps for monitoring millions of assets across cloud, on-prem and edge environments.

You will work with data scientists on ML models for anomaly detection and root cause analysis, and help build scalable, distributed systems with an emphasis on data quality, security and compliance.

Qualifications

  • 12+ years in product development and full stack engineering with 5+ years in observability platforms.
  • Bachelor’s degree in CS/engineering or equivalent experience.
  • Experience with Prometheus, Victoria Metrics, Grafana, OpenTelemetry, etc.
  • Hands-on with AIOps tools like BigPanda, PagerDuty, Datadog.
  • Experience with Kubernetes, Docker, microservices, and streaming with Kafka/NATS.
  • Proficient in Go, Python, Java, or C#.

Responsibilities

  • Lead design, development, and deployment of AIOps and Observability platforms.
  • Define technical vision and roadmap for observability initiatives.
  • Collaborate with teams to meet observability needs.
  • Set standards and evaluate new observability technologies.
  • Provide peer reviews on performance, scalability, security, and correctness.
  • Develop AI agents and AI-native tools for faster issue resolution.

Skills

Observability platforms
Kubernetes
Go
Python
Distributed systems
Cloud-native
Prometheus

Education

Bachelor's degree in CS/Engineering

Tools

Prometheus
Victoria Metrics
Grafana
OpenTelemetry
Kafka
NATS
Docker

Job description

NVIDIA is seeking a highly skilled Senior Software Engineer to design and develop AIOps and Observability platforms. You will mentor engineers and collaborate with product managers to define strategy and roadmaps for monitoring millions of assets across cloud, on-prem and edge environments.

You will work with data scientists on ML models for anomaly detection and root cause analysis, and help build scalable, distributed systems with an emphasis on data quality, security and compliance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AIOps & Observability Architect
Senior AIOps & Observability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior AIOps & Observability Platform Architect
Senior AIOps & Observability Platform Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior AI Observability Platform Engineer
Senior AI Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA AI • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior AIOps SRE for AI Data Center Platform
Senior AIOps SRE for AI Data Center Platform

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 276,000
Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000