Senior Observability Engineer: Telemetry for GPU Cloud

Submer - Datacenters That Make Sense

United States

Remote

USD 104,000 - 174,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid-friendly approach
International team diversity
Flexible work environment

Job summary

Radian Arc is seeking a senior engineer to design and build an observability platform for its GPU cloud infrastructure and edge deployments. You will implement low-latency telemetry pipelines, ingesting metrics, logs, and traces across datacenters and edge locations to enable real-time visibility.

You will lead major observability initiatives, standardize telemetry across services, and collaborate with compute, storage, networking, and platform teams to ensure reliable performance and actionable

Qualifications

  • Proven experience operating large distributed infrastructure platforms.
  • Strong background in observability systems and telemetry pipelines.
  • Experience building metrics, logging, tracing, alerting, and dashboards at production scale.
  • Strong programming skills in Go, Python, or Rust.

Responsibilities

  • Design scalable telemetry pipelines for metrics, logs, and traces across distributed GPU infrastructure.
  • Architect observability systems ingesting high-cardinality telemetry from thousands of nodes and services.
  • Build and operate telemetry storage optimized for time-series and event data.
  • Contribute to observability standards across services, including metrics, tracing, logging, and SLOs.
  • Instrument compute, storage, and networking layers to provide visibility into performance bottlenecks.
  • Develop dashboards and tools to expose system health to internal teams and customers.

Skills

Go
Python
Rust

Tools

Prometheus
OpenTelemetry
Grafana
ClickHouse

Job description

Radian Arc is seeking a senior engineer to design and build an observability platform for its GPU cloud infrastructure and edge deployments. You will implement low-latency telemetry pipelines, ingesting metrics, logs, and traces across datacenters and edge locations to enable real-time visibility.

You will lead major observability initiatives, standardize telemetry across services, and collaborate with compute, storage, networking, and platform teams to ensure reliable performance and actionable

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability & Telemetry Platform Engineer
Senior Observability & Telemetry Platform Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior GPU Cloud Deployment & Operations Engineer
Senior GPU Cloud Deployment & Operations Engineer

Submer - Datacenters That Make Sense • United States

Remote
USD 120,000 - 180,000
Hybrid-friendly culture
Competitive compensation
International diversity
Senior AI Infrastructure Engineer Observability & Automation
Senior AI Infrastructure Engineer Observability & Automation

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 356,000
Equity
Benefits
Senior Network Observability Engineer for GPU Cloud
Senior Network Observability Engineer for GPU Cloud

Socket.dev • Sunnyvale (CA), New York (NY)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Staff AI Telemetry & Observability Engineer
Staff AI Telemetry & Observability Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

NVIDIA Corporation • Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Engineer — Telemetry & Observability
Senior AI Infra Engineer — Telemetry & Observability

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000