Senior AI Infra Engineer — Telemetry & Observability

Socket.dev

North Carolina

Hybrid

USD 184,000 - 357,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking an AI Infrastructure Engineer to design and operate telemetry, incident-management automation, and infrastructure catalogs across on-premises and cloud environments. You’ll help standardize data collection, build AI-assisted tooling, and improve observability for critical AI platforms.

The role emphasizes cross-functional leadership, production-grade software principles, and working across engineering, product, and security teams to reduce toil and accelerate incident response.

Qualifications

  • BS degree in CS/CE or related field, or equivalent experience.
  • 8+ years of experience in infrastructure security, platform engineering, or security tooling.
  • Proficiency in Python, Go, Typescript, or Java.
  • Strong understanding of software and infrastructure principles in production.
  • Ability to lead cross-functional initiatives across engineering, product, finance, and security.

Responsibilities

  • Build and operate scalable telemetry pipelines for metrics, logs, traces, and events across on-premise, CSP, and NCP clusters.
  • Establish common instrumentation, collection, storage, and access patterns for telemetry.
  • Deliver dashboards, alerting, and analysis capabilities to improve service visibility and troubleshooting.
  • Standardize and automate incident, maintenance, service on-call workflows across HWInf.
  • Integrate operational data and lifecycle signals to improve ownership, escalation, and post-incident learning.
  • Build reporting and AI-assisted tooling that reduces manual toil and improves responsiveness.
  • Build and maintain physical hardware and software catalogs as trusted sources of truth for infrastructure inventory.
  • Create data models and pipelines to connect clusters, hardware, services, teams, and workflows.
  • Provide self-service discovery so engineers identify what they operate and who owns it.

Skills

Python
Go
Typescript
Java

Education

BS degree in Computer Science, Computer Engineering, or related field

Job description

NVIDIA is seeking an AI Infrastructure Engineer to design and operate telemetry, incident-management automation, and infrastructure catalogs across on-premises and cloud environments. You’ll help standardize data collection, build AI-assisted tooling, and improve observability for critical AI platforms.

The role emphasizes cross-functional leadership, production-grade software principles, and working across engineering, product, and security teams to reduce toil and accelerate incident response.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

NVIDIA Corporation • California (MO)

Hybrid
USD 184,000 - 357,000
Equity compensation
Benefits
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

NVIDIA Corporation • Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Engineer: Telemetry, Observability & CMDB
Senior AI Infra Engineer: Telemetry, Observability & CMDB

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Engineer Observability & Automation
Senior AI Infrastructure Engineer Observability & Automation

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 356,000
Equity
Benefits
Senior AI Infra Engineer – EDA Platforms & Telemetry
Senior AI Infra Engineer – EDA Platforms & Telemetry

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Observability Engineer
Senior AI Infra Observability Engineer

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Remote work flexibility
+1
Senior AIOps & Observability Architect
Senior AIOps & Observability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Senior AI Factory Observability Architect (Remote, Equity)
Senior AI Factory Observability Architect (Remote, Equity)

NVIDIA Corporation • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity options
Comprehensive benefits package
Senior AIOps & Observability Platform Architect
Senior AIOps & Observability Platform Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits