About Penbrothers
Penbrothers is an HR & remote talent management partner and one of the fastest-growing companies in the Philippines. We provide talented Filipinos with global opportunities in high-growth startups and dynamic companies, from the comfort of their own homes.
About the Role
We are looking for a Data Engineer to design and build our data infrastructure on Google Cloud Platform, following a Medallion (Bronze, Silver, Gold) architecture. This role will own the pipelines, transformation logic, and governance layer that turn raw data from our core business systems into clean, trusted, analytics-ready datasets — and will be expected to work at a modern pace by using AI coding tools as a core part of the development workflow.
What You'll Do:
- Design, build, and maintain scalable ELT/ETL pipelines across the Bronze (raw), Silver (cleansed/conformed), and Gold (business-ready) layers of our BigQuery-based Medallion architecture.
- Own ingestion patterns from core business systems (e.g., CRM, ATS, helpdesk, ERP/finance platforms) into the Bronze layer, using batch loads and, where needed, streaming ingestion (Cloud Storage, BigQuery, APIs, etc.)
- Build, test, and maintain dbt models (staging → intermediate → marts) that transform Bronze data into governed Silver and Gold datasets, with documentation and automated tests for every model.
- Implement data quality controls: schema drift detection, deduplication, null/referential checks, and quarantine/error-handling routines across all layers.
- Use Dataplex / Data Catalog to classify, tag, and govern datasets, and to maintain data lineage and metadata documentation for audit and compliance purposes.
- Write and maintain Python-based pipeline code (Cloud Functions, Cloud Run jobs, Airflow/Cloud Composer DAGs), using AI coding assistants to accelerate development, while still applying sound engineering judgment, testing, and peer review to anything AI-generated.
- Orchestrate, schedule, and monitor pipeline dependencies and SLAs, and respond to pipeline failures or data incidents.
- Partner with BI/Analytics and business stakeholders to define and design the Gold-layer data models that power dashboards, KPIs, and self-serve reporting.
- Build and maintain CI/CD pipelines for data workflows (Cloud Build / GitHub Actions) and manage BigQuery datasets, IAM, and infrastructure using version-controlled, infrastructure-as-code practices (e.g., Terraform).
- Monitor and optimize BigQuery cost and performance (partitioning, clustering, slot/reservation usage, query tuning).
- Apply data governance, access control, and data-boundary safeguards appropriate to a multi-client environment where data segregation and compliance matter.
- Document data models, pipeline architecture, and operational runbooks; contribute to team engineering standards and code review practices.
What You Bring:
- 2+ years of hands-on experience in data engineering, building and operating production data pipelines.
- Strong proficiency in SQL and Python, with demonstrated experience designing ETL/ELT pipelines.
- Demonstrated, active use of AI coding assistants (e.g., Claude, GitHub Copilot, Cursor) for writing, debugging, and optimizing Python and SQL code — able to prompt effectively and critically review AI-generated output rather than accepting it blindly.
- Solid grounding in data modeling concepts: dimensional modeling, star schema design, slowly changing dimensions, etc.
- Knowledge of Data Visualization tools like Looker, Looker Studio, Power BI, or a similar Data Viz tool.
- Strong analytical and troubleshooting skills; comfortable resolving ambiguity across multiple source systems.
- Strong English communication skills, both written and verbal, for cross-functional collaboration.
Nice to have:
- Hands-on experience with the Google Cloud Platform data stack: BigQuery, Cloud Storage, Dataplex/Data Catalog, Pub/Sub, Cloud Composer (Airflow), and Cloud Functions or Cloud Run.
- Experience integrating SaaS platforms (e.g., HubSpot, NetSuite, Zoho, Salesforce) into a cloud data warehouse.
- Experience with version control and CI/CD for data workflows (Git, BitBucket, or Cloud Build).
- Experience with data quality/testing tooling (dbt tests, GCP Dataplex, or similar).
- Direct experience implementing or working within a Medallion (Bronze/Silver/Gold) or comparable layered data architecture.
- Practical, production experience with dbt (Core or Cloud) — model layering, testing, documentation, and Git-based workflows.
- Working knowledge of data governance, access control, and metadata/lineage management.
- Exposure to streaming/event-driven data pipelines (Pub/Sub, Dataflow).
- Terraform or other infrastructure-as-code experience.
- Google Cloud certification (Professional Data Engineer or equivalent).
Our Hiring Process & AI Disclaimer:
We use AI tools to streamline our application process—including an initial conversation with an AI Interviewer. All hiring decisions are made entirely by humans: our Talent Acquisition team guides your full candidate journey and makes all evaluation decisions.
You are welcome to use AI for CV refinement and research, but real-time AI assistance during interviews or skill assessments is strictly prohibited to ensure an authentic evaluation.
Read our full AI Disclaimer to learn more
What You’ll Get
Meaningful work & Growth: We take every opportunity to stretch ourselves and deliver an excellent client experience.
Employee as our biggest asset: We are genuinely invested in our people's career and welfare.
Global reach & local impact: Get to work with high-growth startups and dynamic companies from the comfort of your own home.
Powering global startups: We've created 1,400 Filipino jobs that empower global start-ups to focus on growth.