We are seeking an accomplished Senior GCP Data Engineer with 8+ years of overall data engineering experience and strong hands-on expertise in Google Cloud Platform, Python, PySpark, and SQL. The successful candidate will design, build, optimize, and operate scalable batch and streaming data platforms that transform complex data into reliable, secure, and analytics-ready assets. This role requires strong technical leadership, production engineering discipline, and the ability to translate business requirements into resilient cloud data solutions.
Key Responsibilities
- Architect, develop, and maintain scalable data pipelines and data platforms on GCP for batch and near-real-time workloads.
- Build robust ETL and ELT solutions using Python, PySpark, SQL, BigQuery, Cloud Storage, Dataproc, Dataflow, Pub/Sub, and Cloud Composer.
- Design dimensional, normalized, denormalized, and lakehouse data models aligned with analytical and operational requirements.
- Develop advanced SQL transformations, stored procedures, views, and reusable data-processing components.
- Optimize Spark jobs, BigQuery queries, storage layouts, partitioning, clustering, and resource utilization for performance and cost efficiency.
- Implement data ingestion from databases, APIs, files, event streams, and enterprise applications while supporting schema evolution and incremental processing.
- Establish automated data-quality controls, reconciliation, lineage, metadata management, observability, alerting, and operational support.
- Apply security and governance standards through IAM, service accounts, encryption, secrets management, audit logging, and least-privilege access.
- Develop unit, integration, regression, performance, and data-validation tests and embed them within CI/CD pipelines.
- Use infrastructure-as-code and automated deployment practices to deliver consistent, recoverable environments.
- Troubleshoot complex production failures, perform root-cause analysis, and implement permanent corrective and preventive actions.
- Collaborate with architects, analysts, data scientists, application teams, and business stakeholders to convert requirements into maintainable solutions.
- Lead technical design reviews, enforce engineering standards, mentor team members, and contribute to delivery planning and estimation.
Required Qualifications
- 8+ years of professional experience in data engineering, data warehousing, big data, or cloud-based data platform development.
- Strong production experience delivering data solutions on Google Cloud Platform.
- Advanced proficiency in Python for data processing, automation, integration, and reusable framework development.
- Advanced proficiency in PySpark and Apache Spark, including transformations, joins, partitioning, caching, performance tuning, failure recovery, and large-scale distributed processing.
- Expert SQL skills, including complex joins, common table expressions, window functions, query optimization, data modeling, and data quality validation.
- Hands-on experience with BigQuery, Cloud Storage, Dataproc, Dataflow, Pub/Sub, and Cloud Composer or Apache Airflow.
- Strong understanding of ETL/ELT patterns, batch and streaming architectures, data lakes, data warehouses, lakehouses, and orchestration.
- Experience with Git, automated testing, CI/CD, logging, monitoring, incident management, and production support.
- Working knowledge of GCP IAM, networking concepts, encryption, data governance, privacy, and secure data architecture.
- Strong analytical, problem-solving, stakeholder communication, documentation, and technical leadership skills.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related field, or equivalent practical experience.
Key Responsibilities
Job Summary
We are seeking an accomplished Senior GCP Data Engineer with 8+ years of overall data engineering experience and strong hands-on expertise in Google Cloud Platform, Python, PySpark, and SQL. The successful candidate will design, build, optimize, and operate scalable batch and streaming data platforms that transform complex data into reliable, secure, and analytics-ready assets. This role requires strong technical leadership, production engineering discipline, and the ability to translate business requirements into resilient cloud data solutions.
Key Responsibilities
- Architect, develop, and maintain scalable data pipelines and data platforms on GCP for batch and near-real-time workloads.
- Build robust ETL and ELT solutions using Python, PySpark, SQL, BigQuery, Cloud Storage, Dataproc, Dataflow, Pub/Sub, and Cloud Composer.
- Design dimensional, normalized, denormalized, and lakehouse data models aligned with analytical and operational requirements.
- Develop advanced SQL transformations, stored procedures, views, and reusable data-processing components.
- Optimize Spark jobs, BigQuery queries, storage layouts, partitioning, clustering, and resource utilization for performance and cost efficiency.
- Implement data ingestion from databases, APIs, files, event streams, and enterprise applications while supporting schema evolution and incremental processing.
- Establish automated data-quality controls, reconciliation, lineage, metadata management, observability, alerting, and operational support.
- Apply security and governance standards through IAM, service accounts, encryption, secrets management, audit logging, and least-privilege access.
- Develop unit, integration, regression, performance, and data-validation tests and embed them within CI/CD pipelines.
- Use infrastructure-as-code and automated deployment practices to deliver consistent, recoverable environments.
- Troubleshoot complex production failures, perform root-cause analysis, and implement permanent corrective and preventive actions.
- Collaborate with architects, analysts, data scientists, application teams, and business stakeholders to convert requirements into maintainable solutions.
- Lead technical design reviews, enforce engineering standards, mentor team members, and contribute to delivery planning and estimation.
Required Qualifications
- 8+ years of professional experience in data engineering, data warehousing, big data, or cloud-based data platform development.
- Strong production experience delivering data solutions on Google Cloud Platform.
- Advanced proficiency in Python for data processing, automation, integration, and reusable framework development.
- Advanced proficiency in PySpark and Apache Spark, including transformations, joins, partitioning, caching, performance tuning, failure recovery, and large-scale distributed processing.
- Expert SQL skills, including complex joins, common table expressions, window functions, query optimization, data modeling, and data quality validation.
- Hands-on experience with BigQuery, Cloud Storage, Dataproc, Dataflow, Pub/Sub, and Cloud Composer or Apache Airflow.
- Strong understanding of ETL/ELT patterns, batch and streaming architectures, data lakes, data warehouses, lakehouses, and orchestration.
- Experience with Git, automated testing, CI/CD, logging, monitoring, incident management, and production support.
- Working knowledge of GCP IAM, networking concepts, encryption, data governance, privacy, and secure data architecture.
- Strong analytical, problem-solving, stakeholder communication, documentation, and technical leadership skills.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related field, or equivalent practical experience.
At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.
HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.