Overview
We are looking for an AI & Data Engineer who has strong foundational expertise in data engineering and has extended that into building intelligent AI systems. This role sits at the intersection of scalable data infrastructure and production-grade AI — you will be expected to design and build robust pipelines, model data correctly from first principles, and deliver GenAI-powered applications that run reliably in enterprise environments.
You will work across the full data-to-intelligence stack: ingestion, transformation, warehousing, retrieval, reasoning, and deployment. Equal weight is placed on data engineering rigor and AI engineering capability.
Key Responsibilities
Data Engineering
- Design, build, and maintain end-to-end batch and real‑time data pipelines using Python, PySpark, and Databricks across structured, semi-structured, and unstructured data sources.
- Develop and manage ELT/ETL workflows using dbt for transformation and Airflow for orchestration— including testing, documentation, and data quality enforcement.
- Build and maintain cloud data warehouse and lakehouse architectures (Snowflake, BigQuery, Delta Lake) following modelling best practices.
- Implement dimensional models, SCD strategies, and incremental load patterns appropriate to each use case.
- Integrate data from diverse enterprise sources: REST APIs, relational databases, Kafka event streams, and file-based systems.
- Manage data quality, lineage tracking, and metadata documentation as part of standard delivery.
- Optimise query performance, partitioning strategies, and materialization approaches in cloud warehouse environments.
AI & GenAI Engineering
- Architect and build production-grade RAG pipelines combining vector retrieval, graph traversal, and structured data lookups for multi‑hop reasoning over enterprise datasets.
- Build distributed agent‑based systems using LangGraph, MCP, and A2A frameworks with modular services for ingestion, retrieval, and reasoning.
- Design LLM evaluation and validation layers to enforce semantic accuracy and reduce manual review overhead in AI pipelines.
- Develop LLM‑based automation workflows, recommendation systems, and document intelligence solutions.
- Build explainability layers that combine reasoning traces with LLM outputs to meet enterprise compliance and governance requirements.
Infrastructure & Integration
- Deploy data and AI systems on Kubernetes with robust CI/CD pipelines ensuring reproducibility and production reliability.
- Design FastAPI microservices to expose data and AI capabilities to upstream and downstream enterprise systems.
- Apply observability and monitoring best practices across pipelines and AI services—tracking data freshness, pipeline SLAs, and model behaviour.
- Implement caching strategies (Redis) and query optimisation techniques to meet latency requirements at production scale.
Mandatory Experience
Data Engineering
- Strong proficiency in Python for data pipeline development, ETL/ELT automation, and data processing.
- Advanced SQL—window functions, CTEs, query optimisation, and performance tuning on large‑scale datasets.
- Hands‑on experience with dbt for data transformation, modelling, incremental strategies, testing, and documentation.
- Experience with Apache Airflow for orchestrating multi‑step workflows— DAG design, retry logic, dependency management, and SLA monitoring.
- Proficiency with PySpark and Databricks for distributed data processing, including Delta Lake and batch/streaming workloads.
- Experience with cloud data warehouses—Snowflake, BigQuery, or Redshift—including schema design, partitioning, clustering, RBAC, and query optimisation.
- Solid understanding of data modelling paradigms: dimensional modelling (star/snowflake schema), SCDs, and medallion/lakehouse architecture.
- Familiarity with streaming ingestion using Apache Kafka or equivalent event‑driven platforms.
- Experience with data quality frameworks, schema validation, and pipeline observability.
AI & GenAI Engineering
- Experience building production‑grade RAG (Retrieval‑Augmented Generation) pipelines with hybrid retrieval—dense vector search combined with structured or graph‑based retrieval.
- Proficiency with LangChain, LangGraph, or equivalent frameworks for LLM orchestration and agentic pipeline design.
- Working knowledge of vector databases—Pinecone, Qdrant, Weaviate, or equivalent— including embedding model selection and indexing strategies.
- Experience integrating LLM APIs (OpenAI, Azure OpenAI, Anthropic) into production data and AI pipelines.