Data Engineer - Azure Data Lake & PySpark Pipelines

Vantage Data Centers

Greater London

Hybrid

GBP 66,000 - 106,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Vantage Data Centers in London is seeking a Data Engineer to design, build, and operate scalable data pipelines on the Microsoft Azure data platform. You will work with PySpark and SQL to curate datasets that support enterprise reporting and analytics.

You will own pipelines end-to-end, monitor performance, and collaborate with analysts to translate data requirements into practical solutions, while adhering to governance and security standards in a fast-paced environment.

Qualifications

  • Bachelor's degree or equivalent experience in a related field.
  • 3–5 years of experience in data/analytics engineering.
  • Proficiency in Python for building data pipelines, including PySpark.
  • Proficiency in SQL for querying and transformation.
  • Understanding ETL/ELT pipelines and data integration concepts.
  • Experience analyzing enterprise data sources and business rules.
  • Experience building solutions on Microsoft Azure with Azure Data Factory, Synapse, and Data Lake Gen2.
  • Experience with GitHub/Azure DevOps CI/CD workflows.
  • Knowledge of data modeling basics including fact and dimension tables.
  • Strong communication and collaboration across teams in fast-paced environments.
  • Experience working in Agile environments and using Jira.
  • Ability to travel up to 10% as needed.

Responsibilities

  • Design, build, and maintain scalable data pipelines using Python/PySpark on Azure.
  • Develop and operate batch/incremental pipelines with Azure Data Factory and Data Lake Gen2.
  • Implement SQL- and Spark-based transformations for curated datasets for enterprise reporting.
  • Own data pipelines and datasets with monitoring, troubleshooting, and performance tuning.
  • Work with Azure Synapse to support analytical workloads.
  • Collaborate with business analysts to translate data requirements into solutions.
  • Prepare data for advanced analytics and AI use cases, ensuring quality and documentation.
  • Apply data governance, security, and engineering standards for scalable solutions.
  • Participate in code reviews, discussions, and platform improvement initiatives.
  • Identify data quality issues, pipeline risks, and opportunities, communicating them clearly.
  • Develop and maintain PySpark notebooks/jobs to ingest and transform data.
  • Create/modify Azure Data Factory pipelines for batch/incremental ingestion.
  • Implement Spark transformations to write curated data to ADLS Gen2 with established conventions.
  • Create and maintain SQL views/tables in Synapse to support analytics.
  • Respond to pipeline failures and data validation issues.
  • Perform basic Spark performance tuning within architectural standards.
  • Validate outputs with business partners and address discrepancies.
  • Commit code with Git and adhere to PR reviews and branching standards.
  • Update pipeline/dataset/runbook documentation as changes are made.
  • Execute backlog items within sprint timelines.

Skills

Python programming
SQL
Data modeling
Agile development
Communication skills
Team collaboration
Problem solving

Education

Bachelor's degree in Engineering/related field

Tools

Azure Data Factory
Azure Synapse
Azure Data Lake Storage Gen2
GitHub
Azure DevOps
Jira
PySpark
Git

Job description

Vantage Data Centers in London is seeking a Data Engineer to design, build, and operate scalable data pipelines on the Microsoft Azure data platform. You will work with PySpark and SQL to curate datasets that support enterprise reporting and analytics.

You will own pipelines end-to-end, monitor performance, and collaborate with analysts to translate data requirements into practical solutions, while adhering to governance and security standards in a fast-paced environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

London Data Engineer: Azure Pipelines & PySpark
London Data Engineer: Azure Pipelines & PySpark

Vantage Data Centers • Greater London

Hybrid
GBP 50,000 - 70,000
Flexible work policy
Health and wellness benefits
Hybrid Data Engineer: Azure & PySpark
Hybrid Data Engineer: Azure & PySpark

Vantage Data Centers • Greater London

Hybrid
GBP 55,000 - 75,000
Health benefits
Flexible work policy
Training and development opportunities
Data Engineer (Mid-Level)
Data Engineer (Mid-Level)

Vantage Data Centers • Greater London

Hybrid
GBP 66,000 - 106,000
Data Engineer (Mid‑Level ), Global
Data Engineer (Mid‑Level ), Global

Vantage Data Centers • Greater London

Hybrid
GBP 50,000 - 70,000
Flexible work policy
Health and wellness benefits
Data Engineer - Azure, Databricks & Pipelines
Data Engineer - Azure, Databricks & Pipelines

Independent Vetcare Limited • West of England

Hybrid
GBP 37,000 - 77,000
Healthcare Cash Plan
Cycle to Work scheme
Green Cars salary sacrifice scheme
+2
Data Engineer
Data Engineer

RedRock Resourcing • Birmingham

On-site
GBP 40,000 - 60,000
Data Engineer - Azure, Spark & AI Pipelines (Hybrid Madrid)
Data Engineer - Azure, Spark & AI Pipelines (Hybrid Madrid)

AVEVA • Cambridgeshire and Peterborough

Hybrid
GBP 51,000 - 77,000
Tech Lead Data Engineer - Azure Databricks & PySpark
Tech Lead Data Engineer - Azure Databricks & PySpark

EPAM Systems • Greater London

Hybrid
GBP 90,000 - 130,000
EPAM Employee Stock Purchase Plan (ES&
Private medical insurance
Dental care
+3
Azure Data Engineer — ETL & Analytics Pipelines
Azure Data Engineer — ETL & Analytics Pipelines

Tata Consultancy Services • East Midlands

On-site
GBP 60,000 - 90,000
Pension
Health care
Life assurance
+6
Data Engineer - Azure Databricks, ETL & BI Pipelines
Data Engineer - Azure Databricks, ETL & BI Pipelines

Upp Ltd • Greater London

On-site
GBP 55,000 - 65,000
Discretionary performance bonus
29 days holiday
Private healthcare
+7