Anticipated Contract End Date/Length: Approximately 4.5 months Work Set Up: Hybrid (60% office, 40% remote) Clearance Required: BPSS
Our client in the Information Technology and Services industry is looking for a Platform Architect to enhance and optimise an enterprise Data Platform through hands-on Azure Databricks architecture, engineering, operational support, and platform optimisation. The role will focus on enabling Databricks Serverless capabilities, strengthening FinOps practices and platform governance, and enhancing the Databricks Discovery Zone to support migration and consolidation of POSIT/RStudio workloads.
What you will do:
- Assess existing workloads and select appropriate compute models, including Serverless, Jobs Compute, Interactive Compute, and Classic Clusters, based on workload, SLA, performance, utilisation, and cost.
- Configure and implement Databricks Serverless capabilities across notebooks, jobs, SQL workloads, analytical processing, and data pipelines.
- Develop workload placement standards and guidance, identifying scenarios where Serverless may not provide the most cost-effective solution.
- Implement compute policies, autoscaling, quotas, budget controls, and operational guardrails.
- Monitor platform performance and costs, identifying oversized, underutilised, idle, or inefficient resources.
- Define and implement a practical FinOps operating model covering ownership, accountability, environments, projects, applications, cost centres, and teams.
- Establish mandatory tagging standards and integrate automated validation into CI/CD pipelines.
- Deliver granular cost attribution and reporting across workspaces, projects, applications, workloads, and business teams.
- Configure budgets, spend thresholds, alerts, and usage monitoring to proactively manage platform costs.
- Analyse usage and billing data to identify cost anomalies, inefficient workloads, excessive storage, unnecessary data movement, and underutilised resources.
- Enhance the Databricks Discovery Zone to support migration and modernisation of analytics and data science workloads currently running on POSIT/RStudio.
- Enable secure API and external system integrations, external data ingestion, BI and reporting connectivity, scheduling and orchestration, local IDE-based development, and LLM and AI integration.
- Define reusable onboarding, migration, and delivery patterns that reduce technology sprawl while improving platform security, supportability, and delivery speed.
- Design, build, and optimise scalable data ingestion and transformation solutions using Python, PySpark, SQL, and Delta Lake.
- Implement batch and incremental processing patterns, including CDC, schema evolution, reconciliation, data quality controls, and error handling.
- Develop reusable integration frameworks for REST APIs, SaaS platforms, databases, files, object storage, document repositories, enterprise systems, and external data sources.
- Implement secure authentication, secrets management, and credential handling practices.
- Deliver end-to-end data flows from source ingestion through governed and curated data layers supporting analytics, BI, machine learning, and application consumption.
- Support production platforms and drive continuous optimisation across architecture, engineering, implementation, and operational support.
- Document standards, patterns, operational procedures, and architectural decisions for technical and business stakeholders.