We are looking for Data Engineer at Glasgow, Scotland – 2-3 days per week Onsite
Purpose of the Role
To design, build, and maintain scalable data pipelines, data lakes, and data warehouse solutions on AWS. The role focuses on developing high-performance data engineering solutions using PySpark, Spark, Python, and AWS services, enabling secure, reliable, and efficient data processing and analytics across enterprise platforms.
Key Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines using PySpark, Spark, Python, and AWS services.
- Build and optimize data lakes and data warehouse solutions ensuring data quality, security, and accessibility.
- Develop reusable, production-grade ETL/ELT frameworks and data processing solutions.
- Implement orchestration workflows using AWS Step Functions, Airflow, and other automation tools.
- Develop and maintain cloud infrastructure using AWS CloudFormation.
- Collaborate with business stakeholders to understand requirements and translate them into scalable technical solutions.
- Optimize data processing performance, monitoring, and operational support.
- Implement unit testing, code reviews, and CI/CD best practices using GitLab.
- Support platform modernization and migration initiatives leveraging Spark-based architectures.
- Work closely with Data Scientists and Analytics teams to enable AI/ML use cases.
Required Skills & Experience
- Strong hands-on experience in Data Engineering with delivery of production-grade solutions.
- Expertise in PySpark, Apache Spark, Python, and SQL.
- Strong experience designing and optimizing complex data pipelines and ETL/ELT frameworks.
- Hands-on experience with AWS services including:
- S3
- Glue
- Lambda
- Step Functions
- ECS
- IAM
- KMS
- VPC
- SageMaker (preferred)
- Experience with AWS CloudFormation for Infrastructure as Code.
- Strong understanding of data lakes, data warehouses, and distributed data processing.
- Experience with GitLab, CI/CD, Unit Testing, and DevOps practices.
- Excellent problem-solving skills and ability to work independently.
- Strong stakeholder management and communication skills.
Nice to Have
- Experience with Databricks, Delta Lake, Unity Catalog, and migration projects.
- Knowledge of AI/ML and MLOps frameworks.
- Experience with streaming technologies such as Kafka or Kinesis.
Ideal Candidate:
A hands-on Data Engineer with strong expertise in Spark, PySpark, AWS, and CloudFormation, capable of building scalable enterprise data solutions while driving modernization and cloud transformation initiatives.