At cloudandthings.io, we are an engineering-led consultancy focused on building modern Data and AI Platforms for enterprise clients.
We don't just design architectures; we build real systems that power analytics, decision-making, and AI at scale. Our team acts as a force multiplier, accelerating delivery, improving quality, and unlocking measurable business value. We operate with high standards, strong ownership, and a bias for execution.
Who thrives here
We work with smart people who think from first principles, challenge conventional thinking, and take deep ownership and accountability for what they deliver. We value engineers who don't simply identify problems, but take responsibility for solving them and seeing things through.
If you care about solving meaningful problems and building production-grade systems, you'll fit right in.
Overview
As a Data Engineer, you will design, build, and operate modern, cloud-native data platforms across AWS and Azure, leveraging Databricks, Microsoft Fabric, and AWS native services. You will work across the full data lifecycle: ingestion, transformation, modelling, and serving, enabling real-time analytics, reporting, and AI use cases.
We’re looking for Data Engineers across all levels (Junior to Principal) to join our growing Data, Analytics and AI capability.
Key Responsibilities
While the list below is long, an ideal candidate should have working knowledge and experience covering many of the tools and services. The requirements for each project differ over time, and these skills provide an overview of what may typically be required of a Data Engineer.
Software Engineering Foundations
- Strong grounding in software engineering fundamentals (data structures, algorithms, design patterns).
- Proficiency in Python and SQL (additional languages advantageous).
- Experience with Git, CI/CD pipelines, and modern development practices.
- Familiarity with Terraform or Bicep for Infrastructure as Code, including the Databricks Terraform provider and Fabric deployment automation.
- Comfortable working in Linux-based environments.
- Packaging and promoting data assets as code: Databricks Asset Bundles, Databricks Repos, Fabric Git integration and deployment pipelines.
Data Ingestion and Streaming
- Build scalable ingestion pipelines across hybrid and cloud environments.
- Real-time streaming: AWS Kinesis / MSK (Kafka), Azure Event Hubs / Kafka, Databricks Structured Streaming, Fabric Eventstream and Real-Time Intelligence (Eventhouse / KQL databases, Activator).
- Incremental and file-based ingestion: Databricks Auto Loader, Lakeflow Connect, Fabric Dataflows Gen2 and Copy Job.
- Batch ingestion: AWS DataSync, DMS, Azure Data Factory / Synapse Pipelines / Fabric Data Pipelines.
- Integration via APIs, JDBC/ODBC, and CDC pipelines, including Fabric Mirroring (Azure SQL, Cosmos DB, Snowflake, Databricks Unity Catalog) and Delta Sharing.
Storage, Lakehouse and Fabric
- Design and manage data lakes using: Amazon S3, Azure Data Lake Storage Gen2 (ADLS).
- Implement lakehouse architectures using:
- Databricks — Delta Lake, Unity Catalog (catalogs, schemas, external locations, storage credentials, volumes), managed vs external tables, streaming tables and materialised views, Lakehouse Federation, Delta Sharing.
- Microsoft Fabric — OneLake (including shortcuts to S3/ADLS and other Fabric items), Lakehouse, Warehouse, Fabric SQL Database, workspace and domain design, capacity (F SKU) sizing and CU management.
- Table optimisation and maintenance: partitioning, liquid clustering, Z-ordering, OPTIMIZE / VACUUM, V-Order and statistics in Fabric, small-file management, time travel and retention policies.
- Work with modern data formats and open standards: Delta Lake, Parquet, Avro, JSON, Apache Iceberg / Delta UniForm for cross-engine interoperability.
- Experience with: Relational databases (Postgres, SQL Server, Aurora), NoSQL (DynamoDB, Cosmos DB), Caching (Redis).
Data Processing and Transformation
- Build scalable ETL/ELT pipelines using:
- Databricks — PySpark and Spark SQL, Delta Live Tables / Lakeflow Declarative Pipelines, Databricks Workflows, serverless and job compute, Photon, cluster policies, dbt on Databricks.
- Microsoft Fabric — Fabric Data Engineering (Spark notebooks, Spark job definitions, environments and library management), Fabric Warehouse T-SQL, stored procedures, and Dataflows Gen2.
- Implement medallion architecture (Bronze/Silver/Gold) and dimensional models (star schemas, SCD handling) on the lakehouse.
- Develop reusable, testable, and production-grade data pipelines: modular notebooks and Python packages, parameterisation, unit and integration testing, idempotent and restartable loads.
- Performance and cost tuning of Spark workloads: skew and shuffle management, caching strategies, right-sizing compute, autoscaling.
Analytics and AI Enablement
- Design platforms that support Business Intelligence, advanced analytics, machine learning and AI use cases.
- Work with:
- Databricks — Databricks SQL warehouses, AI/BI dashboards and Genie, MLflow, Unity Catalog model registry, Model Serving, Mosaic AI (Vector Search, agent and RAG patterns), feature engineering in Unity Catalog.
- Microsoft Fabric — semantic models and Direct Lake mode, Power BI report and dataset design, DAX fundamentals, Fabric Data Science (notebooks, MLflow experiments), Fabric Data Agents and Copilot experiences.
- AWS — Amazon Redshift / Athena, SageMaker and Bedrock where relevant.
- Support feature engineering, data science workflows, and real-time decisioning systems.
- Implement data quality, observability, and lineage frameworks: DLT expectations, Unity Catalog lineage and system tables, Fabric monitoring hub and lineage view, alerting on freshness and volume anomalies.
Security, Governance & Compliance
- Implement secure, enterprise-grade data platforms: AWS IAM / Azure Entra ID (AAD), RBAC, Managed Identities, service principals.
- Governance:
- Databricks Unity Catalog — three-level namespace design, grants and privilege inheritance, row filters and column masks, attribute-based access, audit logs and system tables.
- Microsoft Fabric — workspace roles and item permissions, OneLake security, sensitivity labels, domains, and integration with Microsoft Purview.
- AWS Lake Formation — fine-grained access control over S3-based lakes.
- Networking: VPC / VNets, Private Endpoints, Direct Connect / ExpressRoute, Databricks secure cluster connectivity / Private Link, Fabric private links and managed VNets.
- Encryption: KMS / Key Vault / TLS, customer-managed keys, secret scopes.
Orchestration and Operations
- Build orchestrated pipelines using: Databricks Workflows (job dependencies, retries, notifications), AWS Step Functions / MWAA (Airflow), ADF / Synapse / Fabric Data Pipelines and notebook scheduling.
- Monitoring & observability: cloud-native monitoring tools (CloudWatch, Azure Monitor, Log Analytics), Databricks system tables and query history, Fabric Capacity Metrics app and monitoring hub.
- Apply best practices across reliability, performance optimisation, and cost optimisation (FinOps): Databricks compute policies, serverless vs classic trade-offs, tagging and chargeback, Fabric capacity smoothing, bursting, and pause/resume strategies.
Requirements
- Bachelor's degree in Engineering, Computer Science, or related field.
- Proven track record of designing and implementing data solutions.
- Knowledge of and experience with Azure and/or AWS Cloud infrastructure and services.
- Hands‑on delivery experience with at least one lakehouse platform (Databricks or Microsoft Fabric), and an appetite to build depth in the other.
- Certifications, such as:
- Microsoft — Fabric Data Engineer Associate (DP-700), Fabric Analytics Engineer Associate (DP-600), Azure Data Engineer Associate (DP-203), Azure Solutions Architect Expert (AZ-305).
- AWS — Data Engineer Associate, Solutions Architect Associate / Professional.
- Any other data-related experience, e.g. working with Hadoop, databases, analytics software, etc.
- Experience with Docker/Containers/Kubernetes/CICD pipelines for data.
- Knowledge of data security and compliance standards.
- Willingness to learn and expand knowledge related to Cloud and Data Technologies.
- Strong problem-solving and analytical skills.
- Self-organising with the ability to prioritise and manage multiple tasks simultaneously.
- Excellent verbal and written communication skills.
- Ability to work collaboratively with clients and team members.
- Willingness to travel to clients as and when required.
What We Offer
- A culture of engineering and an environment where ideas are heard and builders can build.
- Competitive compensation and bonus structure.
- A flexible and supportive work environment that values diversity, work-life balance, and personal growth.
- Opportunities for career advancement and ongoing professional development.
- Ongoing learning and development opportunities to enhance your skills.
- Engaging with cutting-edge technologies and awesome client projects.
- Access to a talented team of professionals and mentors.