Senior Observability Engineer - Scalable AI Inference

Engg

Sunnyvale (CA)

On-site

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cerebras Systems in Sunnyvale, CA, is seeking a Software Engineer focused on Observability to build scalable telemetry, dashboards, and tooling for production AI services. You will design and implement metrics, logging, tracing, and alerting pipelines, define SLIs/SLOs, and collaborate with backend and platform teams to keep systems reliable at a global scale.

You will also mentor on-call practices, contribute to internal developer platforms, and help balance telemetry signal with performance

Qualifications

  • Backend or systems software experience required.
  • Experience with observability tooling and platforms.
  • Experience building and operating telemetry pipelines at scale.

Responsibilities

  • Design and implement observability instrumentation across services and platforms.
  • Build and maintain telemetry pipelines for metrics, logs, and traces at scale.
  • Develop internal observability platforms, libraries, and tooling.
  • Define and operationalize SLIs, SLOs, and alerting strategies.
  • Partner with engineers to make systems debuggable by design.
  • Reduce MTTR by enabling fast root-cause analysis during incidents.
  • Create clear dashboards and alerts reflecting real system health.
  • Balance telemetry signal vs cost, noise, and performance impact.
  • Improve the developer experience around observability and debugging.

Skills

Backend engineering
Distributed systems
Go
C++
Rust
Java
Python

Tools

OpenTelemetry
Prometheus
Grafana
Datadog
Elastic
Jaeger
Tempo

Job description

Cerebras Systems in Sunnyvale, CA, is seeking a Software Engineer focused on Observability to build scalable telemetry, dashboards, and tooling for production AI services. You will design and implement metrics, logging, tracing, and alerting pipelines, define SLIs/SLOs, and collaborate with backend and platform teams to keep systems reliable at a global scale.

You will also mentor on-call practices, contribute to internal developer platforms, and help balance telemetry signal with performance

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Observability Engineer - Scale Telemetry
Senior Observability Engineer - Scale Telemetry

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Staff Software Engineer - Observability
Staff Software Engineer - Observability

Engg • Sunnyvale (CA)

On-site
USD 150,000 - 230,000
Staff Software Engineer - Observability
Staff Software Engineer - Observability

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Staff Software Engineer — Real-Time Inference Systems
Staff Software Engineer — Real-Time Inference Systems

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Senior SDET - AI Inference Platform Quality & Automation
Senior SDET - AI Inference Platform Quality & Automation

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
Staff Software Engineer - Real-Time AI Inference Infra
Staff Software Engineer - Real-Time AI Inference Infra

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 110,000 - 140,000
Senior AI Platform Engineer: Telemetry & Observability
Senior AI Platform Engineer: Telemetry & Observability

Cribl • Olympia (WA)

Remote
USD 185,000 - 215,000
health
dental
vision
+7
Director, AI Inference & Model Scaling
Director, AI Inference & Model Scaling

Cerebras Systems • Sunnyvale (CA)

Hybrid
USD 260,000 - 340,000
Senior AI Platform Engineer, Telemetry & Observability
Senior AI Platform Engineer, Telemetry & Observability

Cribl • Springfield (IL)

Remote
USD 185,000 - 215,000
Health insurance
Dental insurance
Vision
+7