Staff Observability Engineer for Scalable AI Inference

Cerebras

Sunnyvale (CA)

On-site

USD 180,000 - 240,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cerebras Systems in Sunnyvale, CA is seeking a Software Engineer focused on Observability to build and evolve the systems that give us deep visibility into large-scale, performance-critical production systems. You’ll design and implement metrics, logging, tracing, and alerting infrastructure that enables fast debugging, high reliability, and confident operation of complex distributed systems.

This role sits at the intersection of platform engineering, distributed systems, and reliability.

Qualifications

  • Proven backend or systems software development experience.
  • Proficiency in Go, C++, Rust, Java, or Python.
  • Solid understanding of distributed systems, networking, and concurrency.
  • Hands-on observability experience with metrics, logs, and traces.
  • Familiarity with OpenTelemetry, Prometheus, Grafana and related tooling.
  • Ability to design high-signal alerts and scalable telemetry pipelines.
  • Experience debugging large-scale production systems is a plus.
  • Hardware-aware observability in AI/ML or accelerator contexts is a plus.

Responsibilities

  • Design and implement observability instrumentation across services and platforms.
  • Build and maintain telemetry pipelines for metrics, logs, and traces at scale.
  • Develop internal observability platforms, libraries, and tooling.
  • Define and operationalize SLIs, SLOs, and alerting strategies.
  • Partner with engineers to make systems debuggable by design.
  • Reduce MTTR by enabling fast root-cause analysis during incidents.
  • Create clear, actionable dashboards and alerts that reflect real system health.
  • Balance telemetry signal vs cost, noise, and performance impact.
  • Improve the developer experience around observability and debugging.

Skills

Backend & systems programming
Programming languages
Distributed systems
Networking fundamentals
Concurrency & performance
Observability tooling
Metrics, logs & tracing
Monitoring & alerting
OpenTelemetry
Prometheus
Grafana
Datadog/Elastic/Jaeger/Tempo
Service-level indicators & objectives
High-signal alerts
Telemetry pipelines
Developer platforms
Root-cause analysis

Education

Bachelor's degree in a relevant field

Tools

OpenTelemetry
Prometheus
Grafana
Datadog
Elastic
Jaeger
Tempo

Job description

Cerebras Systems in Sunnyvale, CA is seeking a Software Engineer focused on Observability to build and evolve the systems that give us deep visibility into large-scale, performance-critical production systems. You’ll design and implement metrics, logging, tracing, and alerting infrastructure that enables fast debugging, high reliability, and confident operation of complex distributed systems.

This role sits at the intersection of platform engineering, distributed systems, and reliability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Observability Engineer - Scalable AI Inference
Senior Observability Engineer - Scalable AI Inference

New Grad 2026 @ Cerebras Systems • Sunnyvale (CA)

On-site
USD 150,000 - 230,000
Staff Software Engineer - Observability
Staff Software Engineer - Observability

New Grad 2026 @ Cerebras Systems • Sunnyvale (CA)

On-site
USD 150,000 - 230,000
Staff Software Engineer - Observability
Staff Software Engineer - Observability

Cerebras • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer — Real-Time Inference Systems
Staff Software Engineer — Real-Time Inference Systems

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Senior Staff Engineer, Scalable AI Inference & Resilience
Senior Staff Engineer, Scalable AI Inference & Resilience

Cerebras Systems • United States

Remote
USD 180,000 - 250,000
Senior AI Inference Reliability Engineer
Senior AI Inference Reliability Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Inclusive work environment
Opportunities for continuous learning
Startup vitality with job stability
Principal Engineer, AI Inference Reliability
Principal Engineer, AI Inference Reliability

Cerebras • United States

On-site
USD 120,000 - 160,000
Inclusive work environment
Opportunities for continuous learning
Startup vitality with job stability
AI Inference Performance Engineer
AI Inference Performance Engineer

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Staff Software Engineer, Inference Cloud
Staff Software Engineer, Inference Cloud

Cerebras • Sunnyvale (CA)

On-site
USD 120,000 - 150,000