Platform Observability Lead

Databricks

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Databricks in the United States seeks an experienced Tech Lead to shape the future of platform observability and proactive monitoring. This high‑impact role leads complex investigations, designs observability solutions, and drives systemic improvements that enhance customer experience and platform stability.

You will develop monitoring, alerting, and end‑to‑end observability workflows, collaborate with cross‑functional teams to resolve incidents, and mentor engineers on reliability patterns.

Qualifications

  • 6+ years of experience in SRE/DevOps or similar roles.
  • Cloud provider experience (AWS/Azure/GCP) with containers.
  • Hands-on monitoring, logging, and alerting tools (ELK, Prometheus, Grafana, PagerDuty).
  • Proficient in Python for production automation.
  • Owns incident lifecycle from detection to post‑mortem analysis.
  • BS/MS/PhD in CS/CE or related engineering field.

Responsibilities

  • Lead platform incident investigation across cross‑functional teams.
  • Design observability solutions and alerting pipelines.
  • Build automation to reduce mean time to detect and resolve incidents.
  • Mentor junior engineers on observability patterns and service health metrics.
  • Participate in on‑call rotation.

Skills

SRE/DevOps experience
Cloud experience (AWS/Azure/GCP)
Container orchestration (Docker/Kubern
Monitoring/Logging/Alerting
Observability design
Python automation
Incident lifecycle ownership

Education

BS/MS/PhD in CS/CE/Engineering

Tools

Docker
Kubernetes
ELK
Prometheus
Grafana
PagerDuty

Job description

Databricks in the United States seeks an experienced Tech Lead to shape the future of platform observability and proactive monitoring. This high‑impact role leads complex investigations, designs observability solutions, and drives systemic improvements that enhance customer experience and platform stability.

You will develop monitoring, alerting, and end‑to‑end observability workflows, collaborate with cross‑functional teams to resolve incidents, and mentor engineers on reliability patterns.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Observability Lead
Senior Platform Observability Lead

Cacheflow • United States

On-site
USD 140,000 - 210,000
Senior Platform Observability Engineer
Senior Platform Observability Engineer

Databricks Inc. • San Francisco (CA)

On-site
USD 104,000 - 199,000
Head of Observability & Governance Platform
Head of Observability & Governance Platform

Databricks • California (MO)

On-site
USD 229,000 - 314,000
Senior Staff Engineer, Observability & Governance Platform
Senior Staff Engineer, Observability & Governance Platform

Databricks • Mountain View (CA)

On-site
USD 229,000 - 314,000
Senior Staff Eng - Observability & Governance
Senior Staff Eng - Observability & Governance

Databricks • Seattle (WA)

On-site
USD 217,000 - 299,000
Annual performance bonus
Equity options
Comprehensive benefits
Sr Platform Monitoring Engineer
Sr Platform Monitoring Engineer

Databricks • United States

On-site
USD 140,000 - 210,000
Senior Staff Engineer, Observability & Governance
Senior Staff Engineer, Observability & Governance

Databricks • Bellevue (WA)

On-site
USD 217,000 - 299,000
Senior Architect - Observability & Governance Platform
Senior Architect - Observability & Governance Platform

Databricks • Washington

On-site
USD 217,000 - 299,000
Sr Platform Monitoring Engineer
Sr Platform Monitoring Engineer

Cacheflow • United States

On-site
USD 140,000 - 210,000
Senior Staff Engineer, Observability & Governance Architecture
Senior Staff Engineer, Observability & Governance Architecture

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 315,000
Comprehensive benefits
Annual performance bonus
Equity options