We are seeking a highly skilled GCP Data Engineer with 5+ years of experience designing and building scalable, production-grade data platforms. This role requires deep expertise in Google Cloud Platform, Python, PySpark, SQL, and distributed data processing across the complete data lifecycle—including ingestion, transformation, orchestration, storage, quality, and serving.
The ideal candidate is a hands-on engineer with strong experience building batch and real-time streaming pipelines, developing reusable Python data frameworks, optimizing large-scale workloads, and designing high-performance data models in BigQuery. The engineer will help establish reliable, secure, observable, and cost-efficient data solutions using modern GCP architecture and engineering practices.
Key Responsibilities
- Design, develop, and maintain enterprise-scale batch and streaming data pipelines using Python, PySpark, Dataflow, Dataproc, Pub/Sub, BigQuery, and Cloud Composer
- Build modular, reusable, and testable Python frameworks for data ingestion, transformation, validation, and pipeline automation
- Develop high-performance data-processing solutions using PySpark, Spark SQL, and Apache Beam
- Integrate data from relational and NoSQL databases, REST APIs, flat files, cloud storage, third-party platforms, and event streams
- Build real-time streaming pipelines using Pub/Sub and Dataflow, including windowing, triggers, watermarking, late-arriving data, deduplication, and delivery semantics
- Design event-driven architectures supporting real-time analytics, operational reporting, and downstream data products
- Create and maintain workflow orchestration using Cloud Composer and Apache Airflow, including reusable operators, dependency management, retries, alerting, and recovery
- Design scalable BigQuery schemas using dimensional, normalized, denormalized, and lakehouse modeling patterns
- Optimize BigQuery workloads through effective partitioning, clustering, materialized views, query tuning, and cost-management strategies
- Tune Spark workloads by resolving data skew, excessive shuffling, inefficient joins, memory constraints, partitioning issues, and small-file problems
- Develop incremental loading, change data capture, and high-volume data-processing solutions
- Implement automated data-quality checks, reconciliation, schema validation, lineage, logging, monitoring, and alerting
- Develop Python-based data services and APIs using frameworks such as FastAPI
- Apply security best practices using IAM, service accounts, encryption, secret management, and least-privilege access
- Build and maintain CI/CD pipelines for data applications using Cloud Build, GitHub Actions, or equivalent tools
- Provision and manage GCP data infrastructure using Terraform
- Collaborate with architects, analysts, data scientists, and business stakeholders to translate requirements into scalable data solutions
- Create technical documentation, data mappings, operational runbooks, and architecture diagrams
- Participate in code reviews and promote Python, SQL, GCP, and data-engineering best practices
Required Qualifications
- 5+ years of hands-on experience in data engineering, including building production-grade data pipelines and platforms
- Strong programming expertise in Python, including object-oriented programming, reusable modules, exception handling, logging, testing, packaging, and performance optimization
- Advanced experience with PySpark, Spark SQL, and Apache Spark
- Strong knowledge of GCP data services, particularly BigQuery, Dataflow, Dataproc, Pub/Sub, Cloud Composer, and Cloud Storage
- Advanced SQL skills, including complex joins, window functions, CTEs, query optimization, and large-scale data transformation
- Deep expertise in BigQuery, including schema design, partitioning, clustering, performance tuning, workload management, and cost optimization
- Hands-on experience building batch and streaming pipelines with Dataflow and Apache Beam using Python
- Strong experience designing and operating Spark workloads on Dataproc
- Experience developing and managing Airflow DAGs using Cloud Composer
- Experience designing Pub/Sub topics, subscriptions, schemas, retry policies, dead-letter queues, and deduplication strategies
- Strong understanding of distributed processing concepts, including partitioning, shuffling, parallelism, fault tolerance, and delivery guarantees
- Solid understanding of Spark internals, including execution plans, DAGs, Catalyst optimization, memory management, and Spark UI troubleshooting
- Experience with dimensional data modeling, including star and snowflake schemas and SCD Type 1 and Type 2
- Experience processing large-scale partitioned datasets and implementing incremental data loads
- Experience with data formats such as Parquet, Avro, ORC, JSON, and Delta
- Strong understanding of streaming concepts, including event time, processing time, windowing, triggers, watermarking, and exactly-once versus at-least-once processing
- Experience implementing data-quality, observability, reconciliation, and production-monitoring frameworks
- Proficiency with Git, automated testing, Agile delivery, and CI/CD practices
Preferred Qualifications
- Experience building Python APIs and microservices using FastAPI or Flask
- Working knowledge of Terraform and infrastructure-as-code practices
- Experience with CDC technologies and patterns
- Familiarity with Dataform or dbt for SQL-based data transformations
- Knowledge of GCP security, networking, IAM, Secret Manager, and data-governance practices