We are looking for an experienced and driven Lead Data Engineer to design, build, and scale our data infrastructure. You will be at the forefront of building robust pipelines that handle user interaction data, content metadata, and streaming analytics, powering business intelligence, real-time dashboards, and AI-driven insights. This is a high-impact role that combines deep technical execution with cross-functional leadership.
The candidate will have responsibilities across the following functions:
Data Infrastructure and Pipeline Development:
- Design, develop, and maintain scalable ETL/ELT pipelines to process large volumes of user interaction data, content metadata, and streaming analytics.
- Build and optimise data warehouses and data lakes supporting both real-time and batch processing requirements.
- Implement data quality monitoring and validation frameworks to ensure accuracy, completeness, and reliability across all data assets.
- Develop automated data ingestion systems from diverse sources: mobile apps, web platforms, and third-party integrations.
Analytics and Reporting Infrastructure:
- Create and maintain data models supporting business intelligence, user analy> cs, and content performance metrics.
- Build self-service analy> cs plaJorms enabling stakeholders to independently access and explore insights.
- Implement real-time dashboards and alerting systems for critical business and product KPIs.
- Support A/B testing frameworks and experimental data analysis requirements.
Data Architecture and Optimisation:
- Collaborate with software engineers to optimise database performance and query efficiency at scale.
- Design data storage solutions that balance cost, performance, and accessibility.
- Implement data governance practices including cataloguing, lineage tracking, and granular access controls.
- Ensure GDPR and data privacy compliance across all data systems and pipelines.
Leadership and Collaboration:
- Work closely with data scientists, product managers, and analysts to translate requirements into reliable data solutions.
- Lead architectural discussions and drive best practices in code quality, documentation, and engineering standards.
- Mentor junior and mid-level data engineers; champion a culture of knowledge-sharing and continuous improvement.
- Participate in code reviews and contribute to the technical roadmap of the data platform.
The core requirements for the job include the following:
Programming and Core Skills:
- Proficiency in Python (PySpark, Pandas, FastAPI) and advanced SQL for data transformation, orchestration, and analysis.
- Proficiency in the Databricks stack, storage/pipeline optimisation, and cost optimisation.
- Strong understanding of data modelling, schema design, and warehousing principles (Kimball, Data Vault, medallion architecture).
- Experience with version control (Git) and CI/CD pipelines tailored for data workflows.
Big Data and Streaming Technologies:
- Hands-on experience with Apache Spark for large-scale distributed data processing.
- Proficiency with Apache Ka^a or equivalent for real-time streaming data pipelines and event-driven architectures.
- Experience with workflow orchestration tools such as Apache Airflow, Prefect, or Dagster.
Cloud Platforms and Databases:
- Strong hands-on experience with AWS S3 EMR, Glue, Lambda, Redshift, and related services.
- Experience in SQL databases (PostgreSQL, MySQL) and NoSQL systems (MongoDB, Cassandra, Redis).
- Experience with modern data warehouse/lakehouse solutions: Databricks, Snowflake, or BigQuery.
Infrastructure and DevOps:
- Proficiency with Docker and Kubernetes for containerising and deploying data applications.
- Familiarity with infrastructure-as-code (Terraform) and GitOps principles for data platform management.
Experience Requirements:
- 5 To7 years of experience in data engineering or related roles, with at least 2 years in a lead or senior capacity.
- Proven track record of building, deploying, and maintaining production-grade data pipelines at scale.
- Demonstrated experience with streaming/realtime analytics and data platform architecture.
Advanced and Preferred Skills:
- Understanding of data mesh architecture and domain-driven data design principles.
- Experience implementing agent AI or LLM-powered automation for data and analytics workflows.
- Familiarity with data observability tools such as Monte Carlo, Great Expectations, or Soda.
- Experience with data privacy frameworks, security implementations, and compliance tooling.
- Exposure to ML feature stores and MLOps pipelines is a strong plus.