Role: Data Engineer
Experience: 3-7 Years
Location: Bellandur, Bangalore
Summary
We are seeking a talented and motivated Data Engineer with 3+ years of experience to join our growing data team. In this role, you will be instrumental in building, maintaining, and optimizing our data ingestion pipelines, ensuring high data quality and reliability. You'll work with a variety of data sources and technologies, with a strong emphasis on Google Cloud Platform (GCP) services, particularly BigQuery, contributing to the foundation of our data‑driven initiatives.
Responsibilities
- Design & develop robust, scalable, and efficient data ingestion pipelines from various sources (e.g., databases, APIs, streaming data, files) into our data lake/warehouse, focusing on ingesting data into BigQuery.
- Implement and maintain data quality checks, validation rules, and monitoring mechanisms to ensure accuracy, completeness, and consistency of data within BigQuery tables. Identify and resolve data anomalies proactively.
- Collaborate with data analysts and data scientists to understand data requirements and design optimal data models for analytics and reporting within BigQuery (e.g., partitioned and clustered tables, views, external tables).
- Optimize existing data pipelines and BigQuery queries for performance and cost‑efficiency, leveraging BigQuery features like partitioning, clustering, and query optimization techniques.
- Automate data extraction, transformation, and loading (ETL/ELT) processes, utilizing GCP services such as Cloud Functions, Cloud Dataflow, or Cloud Composer (Apache Airflow) for orchestration.
- Create and maintain comprehensive documentation for data pipelines, data models, and data quality standards.
- Provide support for data‑related issues, debug pipeline failures, and resolve timely, including troubleshooting BigQuery job failures and performance issues.
- Work closely with cross‑functional teams including product, engineering, and business intelligence to understand data needs and deliver effective data solutions.
- Research and evaluate new data technologies and tools within the GCP ecosystem to improve our data infrastructure and processes.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related quantitative field.
- 3+ professional years in data engineering or a similar role, focusing on building and maintaining data pipelines.
- Proficiency in Python (highly preferred), Java, or Scala.
- Hands‑on experience with ETL/ELT tools and concepts.
- Proven experience with Google Cloud Platform (GCP) data services, specifically BigQuery.
- Competency in loading data into BigQuery using methods such as batch loading from Cloud Storage or streaming inserts.
- Strong SQL querying skills within BigQuery, including familiarity with its SQL dialect and functions.
- Familiarity with BigQuery table optimizations (partitioning, clustering) and data quality features.
- Experience with data warehousing concepts and data lake architectures.
- Understanding of data quality principles and experience implementing validation techniques.
- Experience with version control systems such as Git.
- Excellent problem‑solving skills and attention to detail.
- Strong communication and interpersonal skills to explain complex technical concepts to non‑technical stakeholders.
Bonus Points
- Experience with other GCP data services such as Cloud Dataflow, Cloud Composer (Apache Airflow), Cloud Storage, Pub/Sub, or Dataproc.
- Familiarity with BigQuery ML or other machine learning concepts.
- Experience with data orchestration tools (e.g., Apache Airflow, Prefect, Dagster).
- Knowledge of data governance and data security best practices within GCP.