Solution Architecture
- Design end-to-end Lakehouse architectures using Databricks
- Define enterprise data architecture aligned with business and technology strategies
- Lead cloud data modernization and migration initiatives
- Develop scalable, secure, and highly available data platforms
Data Engineering
- Design ETL/ELT frameworks using Apache Spark and Databricks
- Architect batch and real-time streaming data pipelines
- Design Medallion Architecture (Bronze, Silver, Gold)
- Define data ingestion strategies for structured, semi-structured, and unstructured data
- Implement Change Data Capture (CDC) and incremental processing
Data Modeling
- Design enterprise data models using dimensional modeling, star and snowflake schemas, and Data Vault
- Build semantic models for reporting and analytics
Databricks Platform Architecture
- Design and optimize Delta Lake, Unity Catalog, Databricks Workflows, Delta Live Tables, Databricks SQL, MLflow, and Lakeflow
- Establish workspace standards and architecture patterns
- Define notebook, job, and code organization strategies
Performance & Cost Optimization
- Optimize Spark jobs and SQL workloads
- Improve partitioning, clustering, and file layouts
- Reduce cloud infrastructure costs through autoscaling and efficient cluster sizing
- Implement monitoring for compute utilization and query performance
Security & Governance
- Implement Unity Catalog for centralized governance
- Design role-based access control (RBAC) and configure row-level and column-level security
- Establish metadata management, lineage, and auditing
- Ensure compliance with enterprise security and regulatory standards
AI & Advanced Analytics
- Architect AI/ML platforms using Databricks
- Enable ML lifecycle management with MLflow
- Design data pipelines for machine learning workloads
- Support Generative AI and Retrieval-Augmented Generation (RAG) solutions by integrating Databricks with vector databases and AI services where appropriate
DevOps & Automation
- Define CI/CD strategies using Azure DevOps or GitHub Actions
- Implement Infrastructure as Code with Terraform (preferred)
- Automate deployment of notebooks, jobs, workflows, and platform configurations
- Establish version control and branching strategies
Leadership & Governance
- Lead architecture reviews and design workshops
- Define technical standards and best practices
- Mentor Data Engineers and Technical Leads
- Collaborate with business stakeholders, project managers, and cloud architects
- Provide technical leadership during project delivery
Required Technical Skills
- Databricks Lakehouse Platform
- Apache Spark (PySpark and Spark SQL)
- Delta Lake, Unity Catalog, Delta Live Tables, Databricks Workflows, Databricks SQL, MLflow
- Structured Streaming
- Cloud Platforms: Microsoft Azure (preferred), Azure Data Lake Storage Gen2, Azure Data Factory, Azure Synapse Analytics, Azure Key Vault, Microsoft Entra ID, AWS or Google Cloud (valuable)
- Programming: Python, PySpark, SQL, Scala (preferred)
- Databases: SQL Server, Oracle, PostgreSQL, Snowflake
- BI & Analytics: Power BI, Tableau
- Semantic Modeling
- DevOps: Git, Azure DevOps, GitHub, Terraform, CI/CD Pipelines
Preferred Experience
- 10–15+ years of IT experience
- 5+ years of hands‑on Databricks architecture and implementation
- Experience designing enterprise data platforms on Azure, AWS, or GCP
- Proven large‑scale data migration and cloud modernization projects
- Experience integrating enterprise systems such as SAP, Oracle, Salesforce, Workday, or other business applications
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field
- Master's degree preferred
Preferred Certifications
- Databricks Certified Data Engineer Professional or Associate
- Databricks Certified Machine Learning Professional (preferred)
- Microsoft Certified: Azure Data Engineer Associate
- Microsoft Certified: Azure Solutions Architect Expert
- Azure Fundamentals (AZ‑900)
Key Competencies
- Enterprise architecture and solution design
- Distributed data processing expertise
- Data governance and security
- Performance tuning and optimization
- Cost optimization
- Stakeholder management
- Leadership and mentoring
- Problem‑solving and analytical thinking
- Communication and presentation skills
- Strategic planning
Typical Project Responsibilities
- Design enterprise Lakehouse architectures
- Migrate legacy data warehouses to Databricks
- Build scalable ETL/ELT pipelines using Spark
- Implement Delta Lake and Unity Catalog
- Establish governance, security, and metadata management
- Optimize performance and cloud costs
- Deliver curated datasets for BI, AI/ML, and analytics
- Define DevOps, CI/CD, and operational best practices
- Lead architecture reviews and guide technical teams throughout the project lifecycle
A Databricks Architect serves as the technical leader for modern data platforms, ensuring that solutions are scalable, secure, governed, and aligned with business objectives while enabling advanced analytics and AI workloads.