Job Description:
Data Engineer
Function: Data & Analytics | GCP
About the Role
We are looking for a skilled and motivated Data Engineer to join our data platform team. In this role, you will design, build, and maintain scalable data pipelines and infrastructure on Google Cloud Platform (GCP). You will work closely with data analysts, data scientists, and platform teams to ensure reliable, efficient, and cost-effective data movement and transformation across the organization
Key Responsibilities
Data Pipeline Development
- Design, develop, and maintain robust batch and streaming data pipelines using Apache Airflow (Astronomer / Cloud Composer) and Cloud Scheduler.
- Build and manage real-time data ingestion pipelines using Confluent Kafka and Google Pub/Sub.
- Develop and optimize data processing jobs using Dataflow (Apache Beam) and Dataproc (PySpark).
Data Storage & Management
- Design and manage data models, tables, and datasets in BigQuery for performance and cost efficiency.
- Manage data storage in Google Cloud Storage (GCS) including partitioning, lifecycle policies, and access controls.
Infrastructure & DevOps
- Provision and manage cloud infrastructure using Terraform (Infrastructure as Code) for platform and ML model deployments.
- Build and maintain CI/CD pipelines for automated testing, deployment, and versioning of data pipelines - including schema and contract validation gates to ensure data integrity across environments.
- Support MLOps pipeline build-out, enabling reliable model training, versioning, and deployment workflows.
- Implement and manage storage tiering automation to optimize data retention, access patterns, and cost across GCS and BigQuery.
- Deploy and manage containerized data services using Cloud Run.
- Manage source code and collaboration via Cloud Repository / GitHub.
- Handle credentials and sensitive configurations securely using Secret Manager.
Streaming & Event-Driven Architecture
- Design and implement event-driven data architectures using Confluent Kafka and Google Pub/Sub.
- Ensure low-latency, high-throughput data delivery across systems.
Collaboration & Data Quality
- Partner with Data Analysts and Data Scientists to understand data needs and deliver reliable datasets.
- Implement data quality checks, monitoring, and alerting across pipelines.
- Document pipeline architecture, data flows, and operational runbooks.
Required Skills & Technologies
Tools & Technologies
Astronomer Apache Airflow, Cloud Composer, Cloud Scheduler
Orchestration
Astronomer Apache Airflow, Cloud Composer, Cloud Scheduler
Batch Processing
Dataproc, PySpark
Stream Processing
Dataflow (Apache Beam), Confluent Kafka, Pub/Sub
Data Warehouse
BigQuery
Storage
Google Cloud Storage (GCS)
Containerization
Cloud Run
Infrastructure as Code
Terraform
CI/CD
CI/CD Pipelines (Cloud Build / GitHub Actions / Jenkins)
Secret Management
Secret Manager
Source Control
Cloud Repository, Git
Programming
Python, PySpark, SQL
MLOps
ML pipeline orchestration, model deployment, storage tiering
Required Qualifications
- Strong hands-on experience with GCP data services (BigQuery, Dataflow, Dataproc, Pub/Sub, GCS).
- Proficiency in Python and SQL for data pipeline development and transformation.
- Experience with PySpark for large-scale distributed data processing.
- Hands-on experience with Apache Airflow (Astronomer or Cloud Composer).
- Working knowledge of Terraform for infrastructure provisioning and ML model deployments.
- Experience building and maintaining CI/CD pipelines including schema and contract validation gates.
- Familiarity with MLOps practices and supporting model deployment pipelines.
Familiarity with Kafka or event-driven streaming architectures.