Job Title: Databricks Data Architect / Lead Databricks Architect
Job Summary
We are looking for a highly experienced Databricks Data Architect to design and implement enterprise-scale Lakehouse and real-time data platforms. The ideal candidate will have strong hands‑on expertise in Databricks, Apache Spark/PySpark, Delta Lake, Unity Catalog, Structured Streaming, Kafka, AWS, and Medallion Architecture.
The candidate should be capable of owning end‑to‑end architecture, defining technical standards, designing scalable data pipelines, implementing governance and security, and optimizing Databricks platforms for performance and cost.
Key Responsibilities
- - Design and own enterprise Databricks Lakehouse architectures.
- - Define and implement Bronze, Silver, and Gold Medallion Architecture.
- - Architect batch and real‑time data ingestion and processing solutions.
- - Develop scalable pipelines using Databricks, PySpark, Spark SQL, Delta Lake, and Auto Loader.
- - Build real‑time streaming solutions using Structured Streaming and Apache/Confluent Kafka.
- - Design CDC‑based architectures using technologies such as Debezium and Kafka.
- - Implement Delta Lake capabilities including MERGE, schema evolution, time travel, OPTIMIZE, Z‑ORDER, VACUUM, and partitioning strategies.
- - Implement enterprise data governance using Databricks Unity Catalog.
- - Define RBAC/ABAC, table/column‑level security, lineage, auditing, and cross‑workspace governance.
- - Design solutions for data quality, observability, security, and regulatory compliance.
- - Optimize Databricks workloads for performance and DBU/cloud cost.
- - Design cluster policies and leverage Serverless, Photon, auto‑scaling, and appropriate compute strategies.
- - Architect Databricks solutions integrated with AWS services such as S3, Glue, Lambda, Redshift, Athena, EMR, Step Functions, IAM, Lake Formation, CloudWatch, and Secrets Manager.
- - Implement infrastructure using Terraform and/or AWS CloudFormation.
- - Establish CI/CD processes for Databricks notebooks, jobs, pipelines, and infrastructure.
- - Produce architecture diagrams, design specifications, technical roadmaps, and architecture decision records.
- - Work with business, engineering, security, DevOps, and executive stakeholders.
- - Mentor Data Engineers and provide architectural and technical leadership.
Required Skills
- - Strong hands‑on experience with:
- - Databricks
- - PySpark
- - Python
- - SQL
- - Unity Catalog
- - Auto Loader
- - Structured Streaming
- - Medallion Architecture
- - AWS
- - S3
- - AWS Glue
- - Redshift
- - IAM
- - CI/CD
- - Terraform / CloudFormation
Preferred Skills
- - Databricks Certified Data Engineer Professional or equivalent certification
- - AWS Solutions Architect certification
- - Kafka Schema Registry
- - MLflow / Model Registry
- - Databricks AI/ML capabilities
- - RAG and Generative AI architectures
- - LangGraph or similar agent orchestration frameworks
- - Data governance and regulatory compliance experience including HIPAA, GDPR, or PCI‑DSS
Ideal Candidate
We are specifically looking for an architect‑level candidate, not only a Databricks developer. The person should have experience making architecture decisions, building enterprise platforms from the ground up, handling high‑volume batch and streaming workloads, implementing governance, and communicating architecture decisions with both technical and business stakeholders.