Responsibilities
- Perform 24×7 monitoring of Databricks clusters, jobs, workflows, repos, and data pipelines.
- Alert monitoring and first-level resolution.
- First level issue troubleshooting/analysis for cluster failures, auto-scaling issues, and job failures.
- Debug Databricks notebook failures and job errors (Spark, SQL, Delta Lake).
- Rerun/retrigger failed jobs per SOP.
- Monitor data ingestion pipelines (streaming & batch).
- Perform daily health checks.
- Prepare incident summary reports and daily operational dashboards.
- Escalate high severity incidents to L3/Platform Engineering as per SLA.
- Handle workspace/user access requests per RBAC policies.
- Identify recurring issues and report to L3/Platform Engineering.
- Analyze driver/executor logs as first level.
Skills
- 2–5+ years of experience in Big Data / Cloud Data Platform Support.
- Hands-on knowledge of Databricks platform (clusters, jobs, repos, MLflow, warehouse).
- Hands on experience in UNIX, SQL, Shell Scripting.
- Hands on experience in Spark UI & job debugging.
- Understanding of CI/CD pipelines (Azure DevOps).
- Understanding of Apache Spark, Azure Cloud.
- SQL, Shell scripting, Unix, Spark UI, Big Data.
Seniority level
Employment type
Job function
Industries
- IT Services and IT Consulting
Referrals increase your chances of interviewing at Optimum Data Analytics by 2x
Sign in to set job alerts for “Application Support Engineer” roles.