Data Engineer

Scrumconnect Consulting

Newcastle upon Tyne

Hybrid

GBP 39,000 - 65,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Scrumconnect is seeking a Data Engineer to design, build and maintain data pipelines on a modern AWS-native stack. You will use Apache Spark and PySpark for distributed processing, Apache Airflow for orchestration, and several AWS services to store and analyze data.

You will also implement infrastructure as code with Terraform and ensure data governance and security. The role requires translating customer needs into scalable, reliable data assets, and collaborating with cross-functional teams in

Qualifications

  • Hands-on data engineering role building data pipelines on AWS cloud.
  • Experience with Spark/PySpark and Airflow for orchestration.
  • Familiarity with Terraform, Docker and GitLab CI/CD practices.

Responsibilities

  • Data pipeline development: Build and maintain scalable data pipelines using Apache Spark and PySpark.
  • Workflow orchestration: Configure and manage Airflow DAGs for reliable scheduling.
  • Root cause analysis: Identify and resolve data quality and pipeline issues.
  • Data modelling: Apply dimensional models and SCD concepts for trusted assets.
  • Infrastructure as code: Provision cloud infra with Terraform; containerise with Docker.
  • Security & governance: Work within IAM and encryption standards in AWS.

Skills

Python
SQL
PySpark
Jupyter Notebooks

Tools

Apache Spark
Apache Airflow
Terraform
Docker
GitLab

Job description

Job Description

Data Engineer

Up to £65k per annum

Apache Spark Python AWS Cloud Data Pipelines

A hands-on data engineering role within a large-scale cloud data programme, responsible for building, maintaining, and troubleshooting data pipelines using Apache Spark, PySpark, Apache Airflow, and a broad suite of AWS services. You will apply strong analytical and engineering skills to deliver trusted, well-governed data assets in a modern, cloud-native environment.

About Scrumconnect

Scrumconnect is a leading UK technology consultancy delivering digital transformation across public and private sectors, contributing to over 20% of the UK's major citizen-facing public services. We specialise in cloud engineering, data platforms, and agile delivery, helping clients build scalable, secure, and user-centred digital solutions that create real impact.

Working Arrangement

This role is hybrid. Candidates must be willing and able to travel to the Newcastle office once per week. Remaining days may be worked remotely from anywhere in the UK.

About The Role

You will work as a Data Engineer on a complex, cloud-based data programme - designing, building, and maintaining data pipelines that process large volumes of data across a modern AWS-native stack. Using Apache Spark and PySpark for distributed data processing, Apache Airflow for orchestration, and a range of AWS services for storage, compute, and analytics, you will help deliver reliable, well-governed data assets to downstream users.

You will apply strong data analysis skills to identify root causes of data issues, work with dimensional data models and slowly changing dimensions, and implement infrastructure as code using Terraform. Familiarity with engineering best practices and the ability to translate customer expectations into applied technical functionality are key to success in this role.

Key responsibilities
  • Data pipeline development Build and maintain scalable data pipelines using Apache Spark and PySpark, processing and transforming large datasets across distributed cloud infrastructure.
  • Workflow orchestration Configure and manage Apache Airflow DAGs for task orchestration, ensuring reliable scheduling, monitoring, and execution of data processing workflows.
  • Root cause analysis Perform data analysis to identify and resolve root causes of pipeline failures and data quality issues - including reviewing EMR output logs and CloudWatch metrics.
  • Data modelling Apply understanding of dimensional data models and slowly changing dimensions (SCD) to design and maintain well-structured, analytically trusted data assets.
  • Infrastructure as code Provision and manage cloud infrastructure using Terraform. Containerise solutions using Docker and manage deployments through GitLab CI/CD pipelines and release tagging.
  • Security & encryption Apply understanding of both Server Side and client-side encryption patterns within AWS. Work within IAM policies and data governance standards appropriate to a regulated government environment.
Technical Skills Required
Languages & Analytics
  • Python - primary language for pipeline development and data processing
  • SQL - used for querying, transformation, and validation across data stores
  • PySpark - for distributed data processing using Apache Spark on AWS EMR
  • Familiarity with basic data structures for constructing robust, scalable solutions
Data processing & orchestration
  • Apache Spark - understanding of distributed data processing architecture and execution
  • Apache Airflow - configuring DAGs and managing task orchestration at scale
  • Jupyter Notebooks - for exploratory data analysis and pipeline prototyping
  • Understanding of dimensional data models and slowly changing dimensions (SCD Types 1, 2, 3)
  • Data analysis skills to identify root cause of issues within pipelines and data assets
AWS services
  • Amazon EMR - running Spark workloads and reviewing output logs
  • Amazon Athena - ad hoc querying of data in S3
  • Amazon Textract and Comprehend - familiarity with AI/ML document extraction and NLP services
  • AWS S3, IAM, CloudWatch, EC2, ECR - core platform services used day-to-day
  • AWS console proficiency - navigating, configuring, and monitoring services
  • Understanding of Server Side and client-side encryption within AWS
Infrastructure, DevOps & delivery
  • Terraform - Infrastructure as Code for provisioning and managing AWS environments
  • Docker - containerisation of data engineering solutions
  • GitLab - source code management, CI/CD pipeline configuration, release tagging, and component versioning
  • Familiarity with engineering best practices
  • Ability to translate customer expectations into applied, functional technical solutions
Technology stack at a glance

PythonPySparkSQLApache SparkApache AirflowJupyter NotebooksDimensional modelling/SCDAWS EMRAmazon AthenaAWS S3AWS IAMAWS CloudWatchAWS EC2/ECRAmazon TextractAmazon ComprehendTerraformDockerGitLab CI/CDGitLab Tags

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Scrumconnect Limited • Newcastle upon Tyne

Hybrid
GBP 70,000 - 95,000
Data Engineer
Data Engineer

慨正橡扯 • Newcastle upon Tyne

On-site
Data Engineer: Real-Time Pipelines & Cloud Data (Hybrid)
Data Engineer: Real-Time Pipelines & Cloud Data (Hybrid)

慨正橡扯 • Newcastle upon Tyne

On-site
Data Engineer - Newcastle
Data Engineer - Newcastle

Accenture UK & Ireland • Newcastle upon Tyne

On-site
GBP 60,000 - 90,000
Hybrid work model
Data Engineer
Data Engineer

Lynx Recruitment • City of Westminster

Hybrid
GBP 54,000 - 66,000
Data Engineer
Data Engineer

Oscar Technology • Greater London

Hybrid
GBP 60,000 - 75,000
Hybrid working (London-based)
Generous benefits package
Ownership in data platform
Data Engineer - AWS, Spark & AI
Data Engineer - AWS, Spark & AI

Datatech Analytics • City Of London

Hybrid
GBP 60,000 - 75,000
Data Engineer
Data Engineer

Square One Resources • Glasgow

Hybrid
Staff Data Engineer | London | Con...
Staff Data Engineer | London | Con...

Wedo Technology Solutions Ltd. • Greater London

On-site
Data Engineer AWS
Data Engineer AWS

SmartSourcing Ltd • City of Edinburgh

Hybrid
GBP 42,000 - 57,000