Experience: 3–8 Years
About Styli Marketplace
Launched in 2019 by Landmark Group, Styli Marketplace is the first e-commerce venture of the group, quickly becoming a leading online destination for fashion and lifestyle across the GCC, including Saudi Arabia, the UAE, Kuwait, Bahrain, and beyond. Styli connects global sellers and creators with millions of fashion-forward customers, offering the latest trends, exceptional value, and convenient services like same-day to 48-hour delivery and flexible payment options. Our mission is to make style accessible, aspirational, and exciting for all, backed by a passionate team fostering a culture of creativity and innovation. At Styli, we aim to revolutionize fashion retail and bring unique experiences to our customers.
Role Overview
We are looking for a skilled and curious Data Engineer to join our Data Platform team. You will design, build, and maintain the pipelines, data models, and infrastructure that turn raw e-commerce data into trusted, analytics-ready assets. This is a hands‑on role with high ownership — you will work closely with data analysts, ML engineers, and product teams to ensure the right data is available at the right time with the right quality.
What You’ll Do
- Design, build, and maintain batch and real-time data pipelines that ingest data from transactional systems, third‑party APIs, event streams, and operational databases.
- Build robust, idempotent, and observable ETL/ELT workflows using Apache Airflow, dbt, or equivalent orchestration and transformation tooling.
- Handle ingestion from diverse sources: MySQL/PostgreSQL, Kafka event streams, S3/GCS data lakes, REST APIs, and SaaS platforms.
- Ensure pipelines are fault‑tolerant, recoverable, and instrumented with data quality checks at every stage.
Data Modelling & Warehouse Engineering
- Design and maintain dimensional and analytical data models (star schema, OBT, wide tables) optimized for BI reporting and ML feature consumption.
- Own the data warehouse layer on BigQuery — including schema design, partitioning strategies, clustering, and cost‑efficient query patterns.
- Write and review dbt models, tests, documentation, and lineage to maintain a well‑governed transformation layer.
- Collaborate with analysts to translate business questions into clean, reusable, and well‑tested data models.
Streaming & Real‑Time Data
- Build and maintain real‑time data pipelines for use cases such as live inventory updates, order event processing, and customer behaviour streams.
- Implement stream processing logic using Apache Flink, Spark Streaming, or ksqlDB for low‑latency data delivery.
- Design Kafka topic schemas, partition strategies, and consumer group management for high‑throughput e‑commerce event flows.
Data Lake & Cloud Infrastructure
- Build and manage a well‑structured data lake on GCS
- Work with cloud‑native data services:
- Collaborate with the Platform/DevOps team on data infrastructure provisioning using Terraform and containerised data workloads on Kubernetes.
Data Quality & Observability
- Implement data quality frameworks using Great Expectations, dbt tests, or Monte Carlo to catch anomalies, nulls, duplicates, and schema drift before they reach downstream consumers.
- Build and maintain data observability dashboards — pipeline SLAs, freshness SLOs, row count anomalies, and schema change alerts.
- Own incident response for data pipeline failures, with clear runbooks and escalation paths.
ML & Analytics Enablement
- Build and maintain feature pipelines for ML models powering personalisation, recommendations, fraud detection, and demand forecasting.
- Partner with ML engineers to design feature stores (Feast or equivalent) and ensure training/serving consistency.
- Enable self‑serve analytics by maintaining clean semantic layers and well‑documented data marts for business intelligence tools (Looker, Metabase, Superset).
What We’re Looking For
Required
- 3–8 years of hands‑on data engineering experience, ideally in a high‑transaction e‑commerce, fintech, or consumer tech environment.
- Strong proficiency in SQL — complex analytical queries, window functions, CTEs, query optimisation, and warehouse‑specific dialects (BigQuery / Redshift).
- Hands‑on experience with Apache Airflow for pipeline orchestration and dbt for transformation layer management.
- Experience building and operating data pipelines at scale — handling millions of events per day with reliability and observability.
- Working knowledge of Apache Kafka or equivalent event streaming platforms for real‑time data ingestion.
- Experience with cloud data warehouses: BigQuery (preferred) and/or AWS Redshift / Athena.
- Proficiency in Python for pipeline development, data transformation, and automation.
- Hands‑on experience with cloud platforms — GCP (BigQuery, Cloud Composer, Dataproc) and/or AWS (Glue, Athena, Kinesis, S3).
- Understanding of data modelling principles: normalisation, dimensional modelling, and analytical patterns.
Good to Have
- Experience with open table formats: Apache Iceberg, Delta Lake, or Apache Hudi.
- Familiarity with Apache Spark (PySpark) for large‑scale distributed data processing.
- Exposure to feature store design and ML pipeline integration.
- Experience with data cataloguing and governance tooling (DataHub, Amundsen, Collibra, or Alation).
- Knowledge of data mesh or data product principles and federated ownership models.
- Experience with real‑time analytics databases: ClickHouse, Druid, or Pinot for high‑throughput OLAP workloads.
- Familiarity with infrastructure‑as‑code (Terraform) and deploying data workloads on Kubernetes.
- Relevant certifications: Google Professional Data Engineer, AWS Data Analytics Specialty, dbt Analytics Engineering.
The Data Problems You’ll Solve at Styli
- Personalisation at scale — building event pipelines that capture every click, view, and purchase to power real‑time recommendations for millions of customers.
- Inventory & demand forecasting — reliable pipelines feeding ML models that optimise stock levels across hundreds of SKUs and multiple markets.
- Campaign & marketing analytics — ingesting data from paid channels, CRM, and app attribution to give growth teams a single source of truth.
- Order & payments intelligence — real‑time event streams for order lifecycle tracking, fraud signals, and settlement reconciliation.
- Cross‑market reporting — unified data models serving multiple country markets with different currencies, catalogues, and fulfilment providers.