Senior Platform Observability Engineer

Databricks Inc.

San Francisco (CA)

On-site

USD 104,000 - 199,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Databricks Inc. is seeking an experienced Tech Lead to shape the future of platform observability and proactive monitoring. You will lead complex investigations, design observability solutions, and drive systemic improvements that enhance customer experience and platform stability.

Responsibilities include developing monitoring solutions, alerting mechanisms, and customer-focused incident detection tools, and collaborating with engineering teams to resolve incidents and perform post-mortems.

Qualifications

  • 6+ years of experience in SRE, DevOps, Production Engineer, or similar.
  • Production-level experience with AWS/Azure/GCP.
  • Experience with Docker and Kubernetes in production.
  • Experience with monitoring/logging/alerting tools (ELK/Prometheus/Grafana/PagerDuty).
  • Strong Python production automation skills.
  • Incident lifecycle ownership from detection to post-mortem.
  • BS/Master/PhD in CS or related engineering field.

Responsibilities

  • Lead platform incident investigations and coordinate cross-functional teams.
  • Design observability solutions and alerting pipelines.
  • Investigate incidents, perform root cause analyses, and drive improvements.
  • Mentor junior engineers on observability patterns and service health metrics.
  • Participate in on-call rotation.

Skills

Python

Education

BS in Computer Science or related field

Tools

Docker
Kubernetes
ELK
Prometheus
Grafana
PagerDuty

Job description

Databricks Inc. is seeking an experienced Tech Lead to shape the future of platform observability and proactive monitoring. You will lead complex investigations, design observability solutions, and drive systemic improvements that enhance customer experience and platform stability.

Responsibilities include developing monitoring solutions, alerting mechanisms, and customer-focused incident detection tools, and collaborating with engineering teams to resolve incidents and perform post-mortems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Observability Lead
Senior Platform Observability Lead

Cacheflow • United States

On-site
USD 140,000 - 210,000
Platform Observability Lead
Platform Observability Lead

Databricks • United States

On-site
USD 140,000 - 210,000
Sr Platform Monitoring Engineer
Sr Platform Monitoring Engineer

Databricks • United States

On-site
USD 140,000 - 210,000
Sr Platform Monitoring Engineer
Sr Platform Monitoring Engineer

Cacheflow • United States

On-site
USD 140,000 - 210,000
Sr Platform Monitoring Engineer
Sr Platform Monitoring Engineer

Databricks Inc. • San Francisco (CA)

On-site
USD 104,000 - 199,000
Head of Observability & Governance Platform
Head of Observability & Governance Platform

Databricks • California (MO)

On-site
USD 229,000 - 314,000
Senior Staff Engineer, Observability & Governance Platform
Senior Staff Engineer, Observability & Governance Platform

Databricks • Mountain View (CA)

On-site
USD 229,000 - 314,000
Senior Architect - Observability & Governance Platform
Senior Architect - Observability & Governance Platform

Databricks • Washington

On-site
USD 217,000 - 299,000
Senior Staff Engineer, Observability & Governance Architecture
Senior Staff Engineer, Observability & Governance Architecture

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 315,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior Staff Eng - Observability & Governance
Senior Staff Eng - Observability & Governance

Databricks • Seattle (WA)

On-site
USD 217,000 - 299,000
Annual performance bonus
Equity options
Comprehensive benefits