Location: UK (Edinburgh)
Lab: IDEA – Innovation, Data Engineering & Artificial Intelligence
Summary
The IDEA Lab is expanding its GCP Data Products capability and is looking for a highly skilled Data Engineer to build, optimise, and scale data pipelines and large‑scale data processing workloads in a cloud‑native environment. You will work with modern distributed data systems, contribute to data platform modernisation, and support large‑scale ingestion, transformation, and analytics workloads across the Business Transaction Banking platform.
Key Responsibilities
Key Responsibilities
- Design and deliver end‑to‑end data pipelines on cloud platforms (GCP preferred).
- Build scalable data ingestion, transformation, and processing workflows using distributed technologies such as Spark, Flink, Storm, or similar.
- Develop robust ELT/ETL processing and migration pipelines, including support for legacy Datastage decommissioning and modernisation.
- Work with a variety of database technologies including relational, NoSQL, MPP and columnar stores (BigQuery, Redshift, Azure SQLDW, HBase, MongoDB).
- Implement streaming and messaging‑based pipelines using Kafka, Pulsar or Pub/Sub.
- Build optimised, scalable data models to support diverse consumption patterns, applying partitioning, sharding, bucketing and aggregation strategies.
- Apply performance tuning and optimisation across storage, compute and query layers.
- Ensure secure handling of data including authentication, authorisation, encryption in transit/at rest, and cloud‑native security controls.
- Implement monitoring, alerting and observability for large‑scale distributed data workloads.
- Use orchestration tools such as Cloud Composer, Airflow or equivalent to operationalise pipelines.
- Contribute to CI/CD, containerisation, Kubernetes‑based deployments, and automated testing practices.
- Participate in data governance, metadata, catalogue and lineage processes as needed.
- Collaborate with engineers, architects and SMEs to deliver stable, high‑quality data products.
Skill Requirements
Required Skills & Experience
- Strong programming skills in Java (preferred), Python, or Scala.
- Hands‑on experience with cloud data services (GCP preferred; Azure/AWS acceptable).
- Practical experience with distributed data processing frameworks such as Spark (Core/SQL/Streaming), Flink, or Storm.
- Strong knowledge of data ingestion, transformation and messaging systems: Kafka, Pulsar, Pub/Sub, etc.
- Understanding of designing scalable data models for varied access patterns.
- Experience with performance tuning, cost‑optimisation and scaling strategies.
- Experience delivering large‑scale big data solutions in batch and/or streaming environments, on cloud or on‑premise.
- Good familiarity with the wider data ecosystem and open‑source frameworks.
- Experience with orchestration (Airflow/Composer) and workflow automation.
- Understanding of DevOps for data systems: CI/CD, containers, Kubernetes and automated testing.
- Knowledge of security for big‑data systems including IAM, encryption, and cluster‑level controls.
- Basic understanding of monitoring and alerting for distributed systems.
- Knowledge of dimensional modelling (star, snowflake, normalized/denormalized).
- Awareness of data governance, cataloguing and lineage tools.
At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.
HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry‑leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.