Senior AIOps & Observability Architect

NVIDIA

Santa Clara (CA)

On-site

USD 200,000 - 322,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

NVIDIA is seeking a Senior Software Engineer to design and develop AIOps and Observability platforms used by internal teams across cloud, on‑prem, data centers and edge. You will mentor engineers, define the observability strategy and roadmap, and collaborate with product managers to drive best practices and scalable solutions.

You will work with data scientists on ML models for anomaly detection and root‑cause analysis, building robust, secure systems; equity and benefits accompany the role.

Qualifications

  • Bachelor’s degree in computer science or related field, or equivalent experience.
  • 12+ years in product development and full stack engineering, with 5+ years in observability platforms.
  • Strong knowledge of observability tools such as Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry, etc.
  • Hands-on knowledge in AIOps tools such as BigPanda, PagerDuty, Datadog, etc.
  • Experience with Kubernetes, Nomad, Docker, and microservices architectures; streaming services to ingest billions of events using NATS, Kafka, etc.
  • Proficient in Go, Python, Java, C#, etc.

Responsibilities

  • Lead the design, development, and deployment of AIOps observability platforms, including metrics, logs, traces, events, alerts, dashboards, and visualizations.
  • Drive the technical vision and roadmap for AIOps and Observability initiatives, aligning with business goals and industry best practices.
  • Collaborate with other teams and customers to understand their observability needs and provide solutions that meet their requirements and expectations.
  • Establish and implement observability standards, guidelines, and processes across NVIDIA; evaluate and adopt new technologies.
  • Provide peer reviews to other engineers including feedback on performance, scalability, security and correctness.
  • Work with data scientists to implement ML models for anomaly detection, forecasting, and root cause analysis on logs, metrics, and events.

Skills

AIOps platforms
Observability tools
Kubernetes
Docker
Microservices
Go
Python
Java
C#
Telemetry pipelines

Education

Bachelor's degree in computer science or related field

Tools

Prometheus
VictoriaMetrics
Vector
Loki
Grafana
Alertmanager
ClickHouse
OpenTelemetry

Job description

NVIDIA is seeking a Senior Software Engineer to design and develop AIOps and Observability platforms used by internal teams across cloud, on‑prem, data centers and edge. You will mentor engineers, define the observability strategy and roadmap, and collaborate with product managers to drive best practices and scalable solutions.

You will work with data scientists on ML models for anomaly detection and root‑cause analysis, building robust, secure systems; equity and benefits accompany the role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AIOps & Observability Platform Engineer
Senior AIOps & Observability Platform Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior AIOps & Observability Platform Architect
Senior AIOps & Observability Platform Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior Software Engineer, AIOps and Observability
Senior Software Engineer, AIOps and Observability

NVIDIA AI • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior AI Observability Platform Engineer
Senior AI Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior AI Factory Observability Architect (Remote, Equity)
Senior AI Factory Observability Architect (Remote, Equity)

NVIDIA Corporation • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity options
Comprehensive benefits package
Remote Senior AI Factory Observability Architect
Remote Senior AI Factory Observability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity options
Comprehensive benefits package
Remote Senior AI Factory Observability Architect
Remote Senior AI Factory Observability Architect

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity options
Health benefits
Senior Observability Platform Engineer & AIOps Lead
Senior Observability Platform Engineer & AIOps Lead

Software Technology Inc. • Buffalo (NY)

On-site
USD 140,000 - 200,000