We are looking for an experienced Data & AI Engineer to join our team and contribute to enterprise and public-sector digital, data, cloud and AI platform engagements.
The role focuses on building production-grade data products, AI-ready data pipelines and software services that ingest, transform, govern and serve both structured and unstructured data.
You will apply strong software-engineering practices across data and AI delivery, including modular development, APIs, automated testing, CI/CD, observability, security and responsible use of AI-assisted coding tools.
Key Responsibilities
- Develop and maintain batch and real-time/streaming data ingestion and transformation pipelines.
- Design and implement lakehouse architectures, curated data models, data products and serving APIs.
- Implement data quality, metadata, lineage, classification and access-control mechanisms.
- Build data pipelines supporting ML/AI use cases, document processing, embeddings, vector search and RAG solutions.
- Develop modular, maintainable and production-ready solutions using Python, SQL and Apache Spark.
- Implement automated testing, code reviews, CI/CD and deployment practices.
- Work with cloud platforms, containers, orchestration and monitoring/observability tools.
- Troubleshoot data and application issues and provide operational support for production workloads.
- Use approved AI coding assistants such as GitHub Copilot, Microsoft Copilot or equivalent enterprise-approved tools to support development, testing, documentation and analysis.
- Independently validate AI-generated code and ensure correctness, security, licensing compliance, performance and maintainability.
- Collaborate with data scientists, software engineers, architects, business stakeholders and delivery teams.
Required Skills & Experience
- 5+ years of professional experience in data engineering, AI engineering, software engineering or a closely related field.
- Strong hands‑on experience with Python and SQL.
- Experience developing ETL/ELT pipelines and data processing solutions.
- Strong knowledge of Apache Spark and modern data/lakehouse architectures.
- Experience with batch and streaming data processing.
- Experience with APIs, Git, automated testing and CI/CD.
- Exposure to cloud data platforms, containers, orchestration and observability.
- Understanding of ML data preparation, embeddings, vector databases/search and Retrieval-Augmented Generation (RAG).
- Understanding of data governance concepts including data quality, lineage, metadata and access control.
- Experience working in complex enterprise or public-sector environments is highly desirable.
- Demonstrated ability to independently review and validate AI-generated code.
Education & Certifications
- Bachelor's degree in Computer Science, Engineering, Information Systems or a related discipline.
Preferred Certifications
- Databricks Certified Data Engineer Associate
- Microsoft Certified: Azure Data Engineer Associate
- AWS Certified Data Engineer – Associate
- Google Cloud Professional Data Engineer
- Microsoft Certified: Azure AI Engineer Associate
Preferred Technologies
Experience with some of the following will be advantageous:
Python | SQL | Apache Spark | Databricks | Azure/AWS/GCP | APIs | Git/GitHub | CI/CD | Docker | Kubernetes | Kafka | Airflow | Lakehouse | Vector Databases | RAG | LLMs | Data Governance | GitHub Copilot | Microsoft Copilot