Data Engineer (Mid-Level)

Vantage Data Centers

Greater London

Hybrid

GBP 66,000 - 106,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Vantage Data Centers in London is seeking a Data Engineer to design, build, and operate scalable data pipelines on the Microsoft Azure data platform. You will work with PySpark and SQL to curate datasets that support enterprise reporting and analytics.

You will own pipelines end-to-end, monitor performance, and collaborate with analysts to translate data requirements into practical solutions, while adhering to governance and security standards in a fast-paced environment.

Qualifications

  • Bachelor's degree or equivalent experience in a related field.
  • 3–5 years of experience in data/analytics engineering.
  • Proficiency in Python for building data pipelines, including PySpark.
  • Proficiency in SQL for querying and transformation.
  • Understanding ETL/ELT pipelines and data integration concepts.
  • Experience analyzing enterprise data sources and business rules.
  • Experience building solutions on Microsoft Azure with Azure Data Factory, Synapse, and Data Lake Gen2.
  • Experience with GitHub/Azure DevOps CI/CD workflows.
  • Knowledge of data modeling basics including fact and dimension tables.
  • Strong communication and collaboration across teams in fast-paced environments.
  • Experience working in Agile environments and using Jira.
  • Ability to travel up to 10% as needed.

Responsibilities

  • Design, build, and maintain scalable data pipelines using Python/PySpark on Azure.
  • Develop and operate batch/incremental pipelines with Azure Data Factory and Data Lake Gen2.
  • Implement SQL- and Spark-based transformations for curated datasets for enterprise reporting.
  • Own data pipelines and datasets with monitoring, troubleshooting, and performance tuning.
  • Work with Azure Synapse to support analytical workloads.
  • Collaborate with business analysts to translate data requirements into solutions.
  • Prepare data for advanced analytics and AI use cases, ensuring quality and documentation.
  • Apply data governance, security, and engineering standards for scalable solutions.
  • Participate in code reviews, discussions, and platform improvement initiatives.
  • Identify data quality issues, pipeline risks, and opportunities, communicating them clearly.
  • Develop and maintain PySpark notebooks/jobs to ingest and transform data.
  • Create/modify Azure Data Factory pipelines for batch/incremental ingestion.
  • Implement Spark transformations to write curated data to ADLS Gen2 with established conventions.
  • Create and maintain SQL views/tables in Synapse to support analytics.
  • Respond to pipeline failures and data validation issues.
  • Perform basic Spark performance tuning within architectural standards.
  • Validate outputs with business partners and address discrepancies.
  • Commit code with Git and adhere to PR reviews and branching standards.
  • Update pipeline/dataset/runbook documentation as changes are made.
  • Execute backlog items within sprint timelines.

Skills

Python programming
SQL
Data modeling
Agile development
Communication skills
Team collaboration
Problem solving

Education

Bachelor's degree in Engineering/related field

Tools

Azure Data Factory
Azure Synapse
Azure Data Lake Storage Gen2
GitHub
Azure DevOps
Jira
PySpark
Git

Job description

Salary: £66,000 - 106,000 per year

Requirements
  • We require a bachelors degree in Engineering, Computer Science, Data Analytics, or a related field, or equivalent experience.
  • We require 3–5 years of experience in data engineering or analytics engineering.
  • We require proficiency in Python for building and maintaining data pipelines, automation, and data processing workflows, including PySpark.
  • We require proficiency in SQL for querying, transformation, and analytical data processing.
  • We require a solid understanding of ETL/ELT pipelines, data transformation patterns, and data integration concepts.
  • We require experience analyzing enterprise data sources to identify data relationships, transformations, and business rules.
  • We require experience building solutions on Microsoft Azure, with exposure to Azure Data Factory, Azure Synapse, Azure Data Lake Storage Gen2, and related analytics services.
  • We require experience working with source control and CI/CD workflows using tools such as GitHub or Azure DevOps.
  • We require working knowledge of data modeling fundamentals, including fact and dimension tables.
  • We require strong communication and interpersonal skills with the ability to collaborate across teams in a fast-paced environment.
  • We require experience working in Agile development environments.
  • We require experience using collaboration and project tracking tools such as Jira or similar tools.
  • We require the ability to travel up to 10% as needed, with the possibility of increased travel as the business evolves.
Responsibilities
  • We design, build, and maintain reliable, scalable data pipelines using Python and PySpark on the Microsoft Azure data platform.
  • We develop and operate batch and incremental data pipelines using Azure Data Factory for orchestration and Azure Data Lake Storage Gen2 as the primary data store.
  • We independently implement SQL- and Spark-based transformations to produce curated datasets that support enterprise reporting, analytics, and downstream consumption.
  • We take ownership of assigned data pipelines and datasets, including monitoring, troubleshooting, and performance optimization in production environments.
  • We work with Azure Synapse, dedicated or serverless where applicable, to support analytical workloads and data consumption patterns.
  • We collaborate with business analysts and cross-functional stakeholders to translate data requirements into practical, working data solutions.
  • We prepare and structure data to support advanced analytics and AI-enabled use cases by ensuring data quality, consistency, and documentation.
  • We apply established data governance, security, and engineering standards to ensure compliant, maintainable, and scalable solutions.
  • We participate in code reviews, technical discussions, and platform improvement initiatives as active contributors.
  • We proactively identify data quality issues, pipeline risks, and improvement opportunities, and communicate them clearly in a fast-paced environment.
  • We develop and maintain PySpark notebooks and jobs to ingest, transform, and curate data within the enterprise data platform.
  • We build and modify Azure Data Factory pipelines for batch and incremental data ingestion.
  • We implement Spark-based transformations that write curated datasets to Azure Data Lake Storage Gen2 using established folder structures and naming conventions.
  • We create and maintain SQL views and tables in Azure Synapse to support analytics and reporting use cases.
  • We respond to pipeline failures, data validation issues, and operational alerts.
  • We perform basic performance tuning of Spark jobs within established architectural patterns and standards.
  • We validate data outputs with business partners and address data defects or discrepancies.
  • We commit code using Git, follow branching standards, and participate in pull request reviews.
  • We update documentation for pipelines, datasets, and operational runbooks as changes are made.
  • We execute assigned backlog items within sprint timelines and raise risks or blockers early.
  • We perform additional duties as assigned by management.
Technologies
  • AI
  • Azure
  • Business Intelligence
  • CI/CD
  • Cloud
  • DevOps
  • ETL
  • Git
  • GitHub
  • Support
  • JIRA
  • Python
  • PySpark
  • SQL
  • Security
  • Serverless
  • Spark
  • Hardware
  • Machine Learning
More

We are Vantage Data Centers, a global provider powering, cooling, protecting, and connecting the technology of hyperscalers, cloud providers, and large enterprises across North America, EMEA, and Asia Pacific. This role is based in our London office with a flexible hybrid schedule of 3 days on site and 2 days from home. You will join our Data Engineering & Business Intelligence team and help build, operate, and scale our enterprise data platform in support of analytics, reporting, and AI-enabled use cases. Our Technology & Systems department drives technological innovation across IT, software development, OT/automation systems, and business process improvement, and we work closely with Sales, Construction, Operations, and Corporate Functions. We value collaboration, speed to value, financial and execution discipline, and a culture of no ego and no arrogance. We offer an above-market total compensation package, comprehensive health and welfare, retirement, and paid leave benefits, along with recognition, training, development, and opportunities to contribute to our company and community.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (Mid‑Level ), Global
Data Engineer (Mid‑Level ), Global

Vantage Data Centers • Greater London

Hybrid
GBP 50,000 - 70,000
Flexible work policy
Health and wellness benefits
Data Engineer (Mid‐Level ), Global
Data Engineer (Mid‐Level ), Global

Vantage Data Centers • Greater London

Hybrid
GBP 55,000 - 75,000
Health benefits
Flexible work policy
Training and development opportunities
Data Engineer - Azure Data Lake & PySpark Pipelines
Data Engineer - Azure Data Lake & PySpark Pipelines

Vantage Data Centers • Greater London

Hybrid
GBP 66,000 - 106,000
Data Engineer
Data Engineer

Independent Vetcare Limited • West of England

Hybrid
GBP 37,000 - 77,000
Healthcare Cash Plan
Cycle to Work scheme
Green Cars salary sacrifice scheme
+2
Data Engineer
Data Engineer

Tria Recruitment • Greater London

On-site
GBP 54,000 - 94,000
Hybrid Data Engineer: Azure & PySpark
Hybrid Data Engineer: Azure & PySpark

Vantage Data Centers • Greater London

Hybrid
GBP 55,000 - 75,000
Health benefits
Flexible work policy
Training and development opportunities
Senior Data Engineer
Senior Data Engineer

Tech4 Ltd • Oxford

Hybrid
GBP 60,000 - 65,000
Hybrid work pattern
Head Office Oxford
London Data Engineer: Azure Pipelines & PySpark
London Data Engineer: Azure Pipelines & PySpark

Vantage Data Centers • Greater London

Hybrid
GBP 50,000 - 70,000
Flexible work policy
Health and wellness benefits
Data Engineer - Tech Lead (Databricks, Pyspark)
Data Engineer - Tech Lead (Databricks, Pyspark)

EPAM Systems, Inc. • Greater London

Hybrid
GBP 90,000 - 130,000
ESPP
Life insurance
Income protection
+11
Data Engineer
Data Engineer

RedRock Resourcing • Birmingham

On-site
GBP 40,000 - 60,000