Senior Observability & Telemetry Engineer - Radian Arc

Jobgether

Schweiz

Vor Ort

CHF 140.000 - 180.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Radian Arc is seeking a Senior Observability & Telemetry Engineer to establish the observability foundation for large-scale GPU cloud and edge infrastructure in Switzerland. You will design and operate telemetry platforms providing real-time visibility across AI workloads, compute, storage, networking, and inference environments.

The role combines observability architecture, infrastructure telemetry, reliability engineering, and customer-facing performance insights, mentoring engineers and

Qualifikationen

  • Proven experience operating observability systems at production scale.
  • Strong programming skills in Go, Python, or Rust.
  • Hands-on experience with Prometheus, OpenTelemetry, Grafana and large-scale telemetry databases like ClickHouse.
  • Experience with GPU cloud/HPC/AI infrastructure and monitoring distributed training or inference workloads.
  • Knowledge of GPU telemetry technologies (NVIDIA DCGM, DCGM Exporter, NVML) and AI workload metrics.
  • Familiarity with cloud-native infra including Kubernetes, automation, CI/CD, and distributed systems.
  • Strong analytical and troubleshooting capabilities to translate telemetry into actionable improvements.
  • Excellent collaboration across infrastructure, networking, storage, compute, and operations teams.
  • Ownership mindset and ability to lead complex observability initiatives.

Aufgaben

  • Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure.
  • Architect and maintain telemetry storage systems for time-series and event data; contribute to instrumentation and SLIs/SLOs.
  • Build dashboards and monitoring tools that provide actionable insights into workload health, GPU utilization, and performance.
  • Collaborate with platform, networking, storage, compute, and operations teams to improve instrumentation and incident response.
  • Provide technical guidance and mentorship to engineers in observability practices; participate in on-call rotations.

Kenntnisse

Go
Python
Rust
Observability
Telemetry
Distributed systems
Analytical
Collaboration

Tools

Prometheus
OpenTelemetry
Grafana
ClickHouse
Kubernetes
NVIDIA DCGM
NVML
gNMI
SNMP

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Observability & Telemetry Engineer - Radian Arc based in Switzerland.


This is a high-impact engineering role focused on building the observability foundation for large-scale GPU cloud and edge infrastructure.


You will design and operate telemetry platforms that provide real-time visibility across distributed AI workloads, compute, storage, networking, and inference environments.


The role combines observability architecture, infrastructure telemetry, reliability engineering, and customer-facing performance insights.


You will work with high-cardinality metrics, logs, traces, and event data across both hyperscale environments and smaller edge deployments.


Your work will directly improve platform reliability, operational efficiency, performance, and incident response.


As a senior engineer, you will lead major initiatives, influence observability standards, and mentor engineers across multiple technical teams.


The environment is international, technically ambitious, and well suited to someone who enjoys solving complex distributed-systems challenges with significant autonomy.


Accountabilities


  • Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure, supporting high-cardinality data from thousands of nodes and services.

  • Architect and maintain telemetry storage systems optimized for large-scale time-series and event data, while contributing to standards for instrumentation, logging, tracing, and SLO implementation.

  • Build comprehensive observability across compute, storage, networking, GPU clusters, inference workloads, and distributed training environments, identifying issues such as GPU throttling, network congestion, storage latency, and hardware degradation.

  • Develop dashboards, monitoring tools, and performance-analysis capabilities that provide internal teams and customers with actionable insights into workload health, GPU utilization, storage throughput, network latency, and inference performance.

  • Build and maintain network and infrastructure telemetry solutions using Python or Go, integrating data from technologies such as NVIDIA Cumulus Linux, VyOS, Citrix NetScaler/WAF, gNMI, SNMP, and streaming telemetry.

  • Develop advanced alerting, anomaly detection, reliability metrics, SLIs, and SLOs, integrating observability signals into operational workflows and incident management processes.

  • Collaborate with platform, networking, storage, compute, and operations teams to improve instrumentation, monitoring, incident response, and platform reliability.

  • Provide technical guidance and mentorship to engineers, promoting effective observability practices and consistent monitoring patterns across the organization.

  • Participate in on‑call rotations supporting production observability and telemetry infrastructure.


Requirements


  • Proven experience operating observability systems and distributed infrastructure platforms at production scale, with strong expertise across metrics, logging, tracing, alerting, dashboards, and telemetry pipelines.

  • Strong programming skills in Go, Python, or Rust, with experience developing telemetry collectors, exporters, automation, or infrastructure tooling.

  • Hands‑on experience with observability technologies such as Prometheus, OpenTelemetry, Grafana, distributed logging platforms, and large‑scale telemetry databases such as ClickHouse or equivalent.

  • Experience working with large‑scale GPU cloud, HPC, or AI infrastructure and monitoring distributed training or inference workloads.

  • Knowledge of GPU telemetry technologies such as NVIDIA DCGM, DCGM Exporter, NVML, GPU Operator telemetry, NVLink, and NVSwitch, as well as AI workload metrics including inference latency, throughput, NCCL health, synchronization latency, and storage I/O.

  • Strong understanding of networking and infrastructure telemetry, including experience with gNMI, SNMP, streaming telemetry, network flow telemetry, RDMA/RoCE, or comparable technologies.

  • Familiarity with cloud‑native infrastructure, including Kubernetes, automation, CI/CD, and distributed systems.

  • Strong analytical and troubleshooting capabilities, with the ability to interpret complex telemetry signals, diagnose performance problems, identify systemic issues, and translate findings into actionable improvements.

  • Excellent collaboration and communication skills, with the ability to work effectively across infrastructure, networking, storage, compute, and operations teams.

  • A proactive, ownership‑oriented mindset and the ability to lead complex observability initiatives in a fast‑moving, technically sophisticated environment.


Benefits


  • Attractive compensation package reflecting your expertise, experience, transferable skills, and market conditions.

  • Permanent, full-time position with an EMEA-based remote work model.

  • Flexible and hybrid‑friendly working environment designed to support international collaboration.

  • Opportunity to work on large‑scale GPU, AI, cloud, networking, and edge infrastructure challenges.

  • Exposure to advanced observability, telemetry, reliability, and distributed‑systems technologies.

  • Opportunity to contribute to major platform initiatives and influence observability standards across multiple engineering teams.

  • Mentorship and leadership opportunities, including the ability to guide engineers and promote best practices.

  • Career growth within a fast‑growing international scale‑up focused on innovative infrastructure solutions.

  • Inclusive and diverse working environment where qualified candidates are considered fairly and supported in their development.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Remote Senior Telemetry & Observability Architect
Remote Senior Telemetry & Observability Architect

Jobgether • Schweiz

Vor Ort
CHF 140.000 - 180.000
Head of Engineering
Head of Engineering

Lyceum • Zürich

Vor Ort
CHF 180.000 - 260.000
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)

Jobgether • Schweiz

Vor Ort
CHF 180.000 - 280.000
Senior Observability Engineer
Senior Observability Engineer

Myjob • Zürich

Vor Ort
CHF 140.000 - 190.000
Staff Storage Platform Engineer (AI Storage) - Radian Arc
Staff Storage Platform Engineer (AI Storage) - Radian Arc

Jobgether • Schweiz

Hybrid
CHF 150.000 - 210.000
Remote-friendly across Europe
Foundational role in AI storage
International and diverse environment
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA Gruppe • Zürich

Vor Ort
CHF 180.000 - 230.000
Senior Software Engineer (Data)
Senior Software Engineer (Data)

Jobgether • Schweiz

Remote
CHF 87.000 - 104.000
Fully remote work environment
Autonomy and ownership
Influence on technical roadmap
+2
Senior Software Engineer – DevOps/Observability Platform
Senior Software Engineer – DevOps/Observability Platform

Nexthink • Lausanne

Hybrid
CHF 52.000 - 88.000
Health insurance
Meal vouchers (11 EUR daily)
Hybrid work model
+6
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Schweiz

Vor Ort
CHF 150.000 - 210.000
Senior Technical Support Engineer, Observe by Snowflake
Senior Technical Support Engineer, Observe by Snowflake

Snowflake • Zürich

Vor Ort
CHF 90.000 - 110.000