Strategic Mandate
As the Data Platform Lead , you are the technical owner and steward of Nielsen's Databricks Lakehouse for Corporate function. You will drive the strategy, architecture, and operational maturity of our Databricks ecosystem end to end - from workspace and Unity Catalog design to cost governance and platform reliability. This role sits at the intersection of platform engineering and data governance: you will define the guardrails, standards, and reusable patterns that let every engineering and analytics team build confidently on a secure, scalable, and cost-efficient foundation. You will act as the "Platform Authority" for Databricks, ensuring the lakehouse scales with the business while remaining governed, observable, and audit-ready. Critically, you will build an AI-driven data platform - architecting AI-native capabilities (feature stores, vector search, GenAI-ready pipelines) directly into the platform, while also using AI-assisted and agentic engineering practices to build, manage, and scale the platform itself faster and more reliably
Core Goals & Responsibilities
- Lakehouse Platform Ownership: Own the end-to-end architecture of the Databricks Lakehouse, including workspace topology, cluster policies, compute strategy (job clusters, SQL warehouses, serverless), and multi-cloud deployment patterns (AWS/Azure/GCP).
- Unity Catalog & Governance: Design and enforce the global data governance model - catalogs, schemas, access controls, lineage, and audit logging - via Unity Catalog, ensuring consistent metadata management and fine-grained access across business units.
- Data Engineering Standards: Define and enforce best practices for Delta Lake table design, Delta Live Tables (DLT) pipelines, medallion architecture (bronze/silver/gold), and performance optimization (Z-ordering, liquid clustering, partitioning, file compaction).
- Platform Reliability & Automation: Partner with CloudOps and DevOps to industrialize the platform through Infrastructure-as-Code (Terraform/Databricks Asset Bundles), CI/CD pipelines, and monitoring, targeting 99.99% platform availability.
- FinOps & Cost Governance: Own cost transparency and optimization across the Databricks estate - cluster right-sizing, workload isolation, chargeback/showback models, and proactive budget alerting - bringing "Strategic Foresight" to cloud spend.
- Cross-Functional Enablement: Work directly with Data Engineering, Analytics, Data Science, and business stakeholders (Finance, HR, Product) to translate requirements into platform capabilities and self-service patterns.
- Governed AI Enablement: Own the governed rollout of AI use cases and native AI features (e.g., Databricks Genie, AI/BI Dashboards, Mosaic AI) on the platform. You hold the keys to unlocking GenAI and natural-language analytics for the business, safely, within Unity Catalog's governance and access-control boundaries.
- Community & Mentorship: Build and lead a community of practice for Databricks developers across the organization - publishing standards, running enablement sessions, and unblocking delivery friction for engineering teams
Expertise & Technology Stack
- Core Engineering: 8-10+ years of deep experience in data engineering, data platform architecture, or distributed systems, with 3+ years specifically architecting and operating Databricks environments at scale.
- Databricks Platform Engineering: Proven hands-on experience building, managing, running, and scaling a production-grade Databricks Data Platform end-to-end - including Unity Catalog, Delta Lake, Delta Live Tables, Databricks Workflows, Databricks SQL, and cluster/compute policy design.
- AI-Native Platform & AI-Assisted Engineering: Experience architecting AI-native platform capabilities (feature stores, vector search/indexes, GenAI-ready pipelines) into the platform, combined with using AI-assisted and agentic engineering practices (e.g., AI coding copilots, automated testing/ops agents) to build, manage, and scale the platform itself.
- Distributed Compute: Deep hands-on expertise with Apache Spark (batch and structured streaming) for large-scale data processing and performance tuning.
- Frameworks & Languages: Expert-level proficiency in Python and SQL, with working knowledge of dbt for transformation-layer standardization.
- Infrastructure & Automation: Proven experience with Terraform (or Databricks Asset Bundles) and CI/CD tooling to codify and automate platform provisioning and pipeline deployment.
- Security & Governance: Strong grasp of data security, RBAC/ABAC, PII handling, and compliance requirements within a governed lakehouse.
- Analytical Thinking: Ability to challenge status-quo assumptions, evaluate architectural trade-offs, and provide a pragmatic pivot when technical or cost risks arise
Requirements & Qualifications
- Experience: 8-10+ years in a Senior or Lead Data Engineering / Data Platform capacity, including direct ownership of a production Databricks environment.
- Education/Certification: Required: Proven